Как управлять сеансами Live API

Gemini Live API обрабатывает непрерывные потоки аудио или текста, которые называются сеансами. Вы можете управлять жизненным циклом сеанса, начиная с первоначального подтверждения подключения и заканчивая корректным завершением.

Ограничения для сеансов

В случае с Live API сеанс – это постоянное подключение, при котором входные и выходные данные непрерывно передаются по каналу связи.

Если сеанс превышает любое из следующих ограничений, подключение будет разорвано. Однако Live API предлагает несколько вариантов действий, которые помогут вам справиться с ограничениями, связанными с сеансами (см. ниже).

  • Окно контекста сеанса ограничено 128 000 токенов.

    Из-за этого ограничения на размер окна контекста максимальная продолжительность сеанса в зависимости от способа ввода будет примерно следующей:

    • Сеансы ввода только аудио ограничены 15 минутами.
    • Продолжительность видео и аудио ограничена двумя минутами.
  • Продолжительность подключения ограничена примерно 10 минутами.

    Вы получите уведомление о том, что связь скоро будет разорвана примерно за 60 секунд до этого.

Вот несколько вариантов решения этой проблемы:

  • Сжать окно контекста сеанса, чтобы сервер автоматически поддерживал размер контекста в пределах лимита.

  • Возобновление сеанса позволяет не терять контекст разговора при кратковременном отключении сети или после получения уведомления Скоро уходим.

Как начать сеанс

Полный фрагмент кода, показывающий, как начать сеанс, можно найти в руководстве по началу работы с Live API.

Обновление в середине сеанса

Модели Live API поддерживают следующие расширенные возможности для обновлений в середине сеанса:

Как добавлять обновления контента

Во время активного сеанса можно добавлять инкрементные обновления. Используйте этот метод, чтобы отправлять текстовые запросы, устанавливать контекст сеанса или восстанавливать его.

  • Если контекст длинный, рекомендуем предоставить краткий пересказ сообщения, чтобы освободить окно контекста для последующих взаимодействий.

  • Для коротких контекстов можно отправлять пошаговые взаимодействия, чтобы представить точную последовательность событий, как в приведенном ниже фрагменте кода.

Swift

// Define initial turns (history/context).
let turns: [ModelContent] = [
  ModelContent(role: "user", parts: [TextPart("What is the capital of France?")]),
  ModelContent(role: "model", parts: [TextPart("Paris")]),
]

// Send history, keeping the conversational turn OPEN (false).
await session.sendContent(turns, turnComplete: false)

// Define the new user query.
let newTurn: [ModelContent] = [
  ModelContent(role: "user", parts: [TextPart("What is the capital of Germany?")]),
]

// Send the final query, CLOSING the turn (true) to trigger the model response.
await session.sendContent(newTurn, turnComplete: true)

Kotlin

// Define initial turns (history/context).
val turns = listOf(
  content("user") {
    text("What is the capital of France?")
  },
  content("model") {
    text("Paris")
  }
)

// Send history, keeping the conversational turn OPEN (false).
turns.forEach {
  session.send(
    content = it,
    turnComplete = false
  )
}

// Define the new user query.
val newTurn = content("user") {
  text("What is the capital of Germany?")
}

// Send the final query, CLOSING the turn (true) to trigger the model response.
session.send(
  content = newTurn,
  turnComplete = true
)

Java

// Define initial turns (history/context).
List turns =
    Arrays.asList(
        new Content.Builder().setRole("user").addText("What is the capital of France?").build(),
        new Content.Builder().setRole("model").addText("Paris").build());

for (Content turn : turns) {
  session.send(
    turn,
    false // turnComplete: false
  );
}

// Define the new user query.
Content newTurn = new Content.Builder().addText("What is the capital of Germany?").build();

// Send the final query, CLOSING the turn (true) to trigger the model response.
session.send(
  newTurn,
  true // isTurnComplete: true
);

Web

const turns = [{ text: "Hello from the user!" }];

await session.send(
  turns,
  false // turnComplete: false
);

console.log("Sent history. Waiting for next input...");

// Define the new user query.
const newTurn [{ text: "And what is the capital of Germany?" }];

// Send the final query, CLOSING the turn (true) to trigger the model response.
await session.send(
    newTurn,
    true // turnComplete: true
);
console.log("Sent final query. Model response expected now.");

Dart

// Define initial turns (history/context).
final List turns = [
  Content(
    "user",
    [Part.text("What is the capital of France?")],
  ),
  Content(
    "model",
    [Part.text("Paris")],
  ),
];

// Send history, keeping the conversational turn OPEN (false).
await session.send(
  input: turns,
  turnComplete: false,
);

// Define the new user query.
final List newTurn = [
  Content(
    "user",
    [Part.text("What is the capital of Germany?")],
  ),
];

// Send the final query, CLOSING the turn (true) to trigger the model response.
await session.send(
  input: newTurn,
  turnComplete: true,
);

Unity

// Define initial turns (history/context).
List turns = new List {
    new ModelContent("user", new ModelContent.TextPart("What is the capital of France?") ),
    new ModelContent("model", new ModelContent.TextPart("Paris") ),
};

// Send history, keeping the conversational turn OPEN (false).
foreach (ModelContent turn in turns)
{
    await session.SendAsync(
        content: turn,
        turnComplete: false
    );
}

// Define the new user query.
ModelContent newTurn = ModelContent.Text("What is the capital of Germany?");

// Send the final query, CLOSING the turn (true) to trigger the model response.
await session.SendAsync(
    content: newTurn,
    turnComplete: true
);

Как изменить системные инструкции во время сеанса

Доступно только при использовании Agent Platform Gemini API (formerly Vertex AI) в качестве поставщика API.

Вы можете изменить системные инструкции во время активного сеанса. Используйте эту функцию, чтобы адаптировать ответы модели, например изменить язык или тон.

Чтобы обновить системные инструкции во время сеанса, отправьте текстовый контент с ролью system. Обновленные инструкции будут действовать до конца сеанса.

Swift

await session.sendContent(
  [ModelContent(
    role: "system",
    parts: [TextPart("new system instruction")]
  )],
  turnComplete: false
)

Kotlin

// In a coroutine scope
session.send(
    content = content("system") {
        text("new system instruction")
    },
    turnComplete = false
)

Java

session.send(
    new Content.Builder()
        .setRole("system")
        .addText("new system instruction")
        .build(),
    /* turnComplete: */ false
);

Web

Not yet supported for Web apps - check back soon!

Dart

try {
  await _session.send(
    input: Content(
      'system',
      [Part.text('new system instruction')],
    ),
    turnComplete: false,
  );
} catch (e) {
  print('Failed to update system instructions: $e');
}

Unity

try
{
    await session.SendAsync(
        content: new ModelContent(
            "system",
            new ModelContent.TextPart("new system instruction")
        ),
        turnComplete: false
    );
}
catch (Exception e)
{
    Debug.LogError($"Failed to update system instructions: {e.Message}");
}

Сжатие окна контекста

Нажмите на поставщика Gemini API, чтобы посмотреть контент и код, относящиеся к нему.

Live API Окно контекста сеанса хранит потоковые данные в реальном времени (25 токенов в секунду для аудио и 258 токенов в секунду для видео), а также другой контент, в том числе текстовые входные данные и выходные данные модели. Для всех моделей Live API действует ограничение на размер окна контекста сеанса – 128 000 токенов.

По умолчанию из-за ограничения на размер окна контекста максимальная продолжительность сеанса в зависимости от способа ввода будет примерно следующей:

  • Сеансы ввода только аудио ограничены 15 минутами.
  • Продолжительность видео и аудио ограничена двумя минутами.

В длительных сеансах по мере развития разговора накапливается история аудио- и/или видеотокенов. Если история превышает лимит модели, она может галлюцинировать, работать медленнее или сеанс может быть принудительно завершен.

Чтобы увеличить продолжительность сеансов, можно включить сжатие контекстного окна, задав поле contextWindowCompression как часть LiveGenerationConfig. Если эта функция включена, сервер использует скользящее окно, чтобы автоматически удалять самые старые реплики или обобщать их, чтобы размер контекста не превышал заданные по умолчанию или указанные вами ограничения. Системные инструкции не удаляются и всегда остаются в начале окна контекста.

С точки зрения пользователя это позволяет теоретически бесконечно продлевать сеанс, поскольку память постоянно управляется.

Вы можете настроить механизм скользящего окна, а также при необходимости указать количество токенов, при котором будет запускаться сжатие (см. доступные настройки и значения ниже). Вот несколько общих рекомендаций по использованию этих настроек:

  • Если установить значение targetTokens слишком низким, для непрерывных потоков будет доступно больше контекста, но модель будет быстро "забывать" более ранние сообщения.

  • Если установить значение targetTokens ближе к triggerTokens, будет сохраняться больше памяти, но сжатие будет выполняться гораздо чаще.

Параметр Значение по умолчанию для скользящего окна, если оно не задано в конфигурации Минимальное значение Максимальное значение
triggerTokens
длина контекста до запуска сжатия
80% от лимита окна контекста модели 5000 128 000
targetTokens
целевое количество токенов, которые нужно сохранить.
50% от значения triggerTokens
  • Если значение triggerTokens не задано явно, то по умолчанию targetTokens составляет 50% от значения по умолчанию triggerTokens.
  • Значение параметра targetTokens должно быть меньше, чем у triggerTokens.
0 128 000

Swift


// ...

let liveModel = FirebaseAI.firebaseAI(backend: .googleAI()).liveModel(
  modelName: "GEMINI_LIVE_API_MODEL_NAME",
  // Enable context window compression.
  // (Optional) Configure the number of tokens in the context window that triggers the compression.
  generationConfig: LiveGenerationConfig(
    responseModalities: [.audio],
    contextWindowCompression: ContextWindowCompressionConfig(
      triggerTokens: 10000,
      slidingWindow: SlidingWindow(
        targetTokens: 2000,
      )
    )
  )
)

Kotlin


// ...

val liveModel = Firebase.ai(backend = GenerativeBackend.googleAI()).liveModel(
    modelName = "GEMINI_LIVE_API_MODEL_NAME",
    // Enable context window compression.
    // (Optional) Configure the number of tokens in the context window that triggers the compression.
    generationConfig = liveGenerationConfig {
        responseModality = ResponseModality.AUDIO,
        contextWindowCompression = ContextWindowCompressionConfig(
            triggerTokens = 10000,
            slidingWindow = SlidingWindow(targetTokens = 2000)
        )
    }
)

Java


// ...

LiveGenerativeModel lm = FirebaseAI.getInstance(GenerativeBackend.googleAI()).liveModel(
        "GEMINI_LIVE_API_MODEL_NAME",
        // Enable context window compression.
        // (Optional) Configure the number of tokens in the context window that triggers the compression.
        new LiveGenerationConfig.Builder()
                .setResponseModality(ResponseModality.AUDIO)
                .setContextWindowCompression(
                        new ContextWindowCompressionConfig(10000, new SlidingWindow(2000))
                )
                .build()
);

Web


// ...

const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });

const liveModel = getLiveGenerativeModel(ai, {
  model: "GEMINI_LIVE_API_MODEL_NAME",
  // Enable context window compression.
  // (Optional) Configure the number of tokens in the context window that triggers the compression.
  generationConfig: {
    responseModalities: [ResponseModality.AUDIO],
    contextWindowCompression: {
      triggerTokens: 10000,
      slidingWindow: {
        targetTokens: 2000,
      },
    },
  },
});

Dart


// ...

final _liveModel = FirebaseAI.googleAI().liveGenerativeModel(
  model: 'GEMINI_LIVE_API_MODEL_NAME',
  // Enable context window compression.
  // (Optional) Configure the number of tokens in the context window that triggers the compression.
  liveGenerationConfig: LiveGenerationConfig(
    responseModalities: [ResponseModalities.audio],
    contextWindowCompression: ContextWindowCompressionConfig(
      triggerTokens: 10000,
      slidingWindow: SlidingWindow(targetTokens: 2000),
    ),
  ),
);

Unity


// ...

var liveModel = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetLiveModel(
    modelName: "GEMINI_LIVE_API_MODEL_NAME",
    // Enable context window compression.
    // (Optional) Configure the number of tokens in the context window that triggers the compression.
    liveGenerationConfig: new LiveGenerationConfig(
        responseModalities: new[] { ResponseModality.Audio },
        contextWindowCompression: new ContextWindowCompressionConfig(
            triggerTokens: 10000,
            slidingWindow: new SlidingWindow(targetTokens: 2000)
        )
    )
);

определять, когда сеанс будет завершен;

Максимальная продолжительность одного непрерывного подключения WebSocket составляет примерно 10 минут. Клиенту отправляется уведомление going away за 60 секунд до завершения подключения. Это позволяет предпринять дальнейшие действия, например возобновить сеанс.

В примере ниже показано, как обнаружить предстоящее отключение, прослушивая уведомление going away:

Swift

for try await response in session.responses {
  switch response.payload {

  case .goingAwayNotice(let goingAwayNotice):
    // Prepare for the session to close soon
    if let timeLeft = goingAwayNotice.timeLeft {
        print("Server going away in \(timeLeft) seconds")
    }
  }
}

Kotlin

for (response in session.responses) {
    when (val message = response.payload) {
        is LiveServerGoAway -> {
            // Prepare for the session to close soon
            val remaining = message.timeLeft
            logger.info("Server going away in $remaining")
        }
    }
}

Java

session.getResponses().forEach(response -> {
    if (response.getPayload() instanceof LiveServerResponse.GoingAwayNotice) {
        LiveServerResponse.GoingAwayNotice notice = (LiveServerResponse.GoingAwayNotice) response.getPayload();
        // Prepare for the session to close soon
        Duration timeLeft = notice.getTimeLeft();
    }
});

Web

for await (const message of session.receive()) {
  switch (message.type) {

  ...
  case "goingAwayNotice":
    console.log("Server going away. Time left:", message.timeLeft);
    break;
  }
}

Dart

Future _handleLiveServerMessage(LiveServerResponse response) async {
  final message = response.message;
  if (message is GoingAwayNotice) {
     // Prepare for the session to close soon
     developer.log('Server going away. Time left: ${message.timeLeft}');
  }
}

Unity

foreach (var response in session.Responses) {
    if (response.Payload is LiveSessionGoingAway notice) {
        // Prepare for the session to close soon
        TimeSpan timeLeft = notice.TimeLeft;
        Debug.Log($"Server going away notice received. Remaining: {timeLeft}");
    }
}

Как возобновить сеанс

Live API поддерживает возобновление сеанса, чтобы не терять контекст разговора. У каждого сеанса есть дескриптор, который можно использовать следующими способами:

  • Поддержание сеанса до достижения лимита времени подключения

    Максимальная продолжительность одного непрерывного подключения WebSocket составляет примерно 10 минут. Вы можете определить, когда подключение будет завершено, прослушивая уведомление going away, а затем продлить сеанс, установив новое подключение с помощью дескриптора сеанса.

  • Возобновление сеанса после потери подключения

    Если подключение прерывается до истечения максимального времени подключения (например, при переходе с Wi-Fi на 5G), сервер сохраняет состояние сеанса в течение примерно 10 минут. В течение этого времени вы можете возобновить сеанс, установив новое подключение с помощью дескриптора сеанса.

  • Возобновление сеанса после длительного периода времени

    После завершения подключения сервер сохраняет состояние сеанса в течение нескольких часов. В течение этого времени вы можете возобновить сеанс, установив новое подключение с помощью дескриптора сеанса. Обратите внимание, что для разных поставщиков Gemini API этот период отличается: для Gemini Developer API он составляет 2 часа, а для Agent Platform Gemini API (formerly Vertex AI) – 24 часа.

По умолчанию возобновление сеанса отключено. Чтобы включить возобновление сеанса, при установке нового подключения передайте пустую конфигурацию возобновления. Если эта функция включена, сервер периодически отправляет обновления, содержащие дескриптор возобновления сеанса. Если сеанс будет отключен, вы сможете восстановить подключение и передать этот дескриптор, чтобы возобновить сеанс с сохраненным контекстом.

Ниже приведены два варианта возобновления сеанса.

Swift

// Local variable to save the active session handle
var activeSessionHandle: String?

// Initialize the session. Passing an empty config requests the server to send SessionResumptionUpdate
var session = try await liveModel.connect(
  sessionResumption: SessionResumptionConfig()
)

// Start receiving responses
for try await message in session.responses {
  // Check for new session handles inside your message handling loop
  switch message.payload {
  case let .sessionResumptionUpdate(updateMessage):
    guard let newHandle = updateMessage.newHandle, updateMessage.resumable else {
      continue
    }
    activeSessionHandle = newHandle
    print("SessionResumptionUpdate: handle \(newHandle)")
  // ... handle other LiveServerMessage types ...
  default:
    break
  }
}

// The following are alternative options to resume a session. Choose only one.

// Option 1: Create and connect a session to resume with the saved handle
if let handle = activeSessionHandle {
  session = try await liveModel.connect(
    sessionResumption: SessionResumptionConfig(handle: handle)
  )
}

// Option 2: Resume the session directly on an existing session object
if let handle = activeSessionHandle {
  try await session.resumeSession(
    sessionResumption: SessionResumptionConfig(handle: handle)
  )
}

Kotlin

// Local variable to save the active session handle
var activeSessionHandle: String? = null

// Initialize the session. Passing an empty config requests the server to send SessionResumptionUpdate
var session = liveModel.connect(
    sessionResumption = SessionResumptionConfig()
)

// Start receiving responses
session.receive().collect { message ->
    // Process other received response types...

    // Check for new session handles inside your message handling loop
    if (message is LiveSessionResumptionUpdate) {
        if (message.resumable == true && message.newHandle != null) {
            activeSessionHandle = message.newHandle
            Log.d("TAG", "SessionResumptionUpdate: handle ${message.newHandle}")
        }
    }
}

// The following are alternative options to resume a session. Choose only one.

// Option 1: Create and connect a session to resume with the saved handle
activeSessionHandle?.let { handle ->
    session = liveModel.connect(
        sessionResumption = SessionResumptionConfig(handle = handle)
    )
}

// Option 2: Resume the session directly on an existing session object
activeSessionHandle?.let { handle ->
    session.resumeSession(
        sessionResumption = SessionResumptionConfig(handle = handle)
    )
}

Java

For Java, session resumption is not yet supported. Check back soon!

Web

// Local variable to save the active session handle
let activeSessionHandle = null;

// Initialize the session. Passing an empty object requests the server to send SessionResumptionUpdate
let session = await liveModel.connect({});

// Start receiving responses
for await (const message of session.receive()) {
  // Process other received response types...

  // Check for new session handles inside your message handling loop
  if (message.type === 'sessionResumptionUpdate') {
    if (message.resumable && message.newHandle) {
      activeSessionHandle = message.newHandle;
      console.log(`SessionResumptionUpdate: handle ${activeSessionHandle}`);
    }
  }
}

// The following are alternative options to resume a session. Choose only one.

// Option 1: Create and connect a session to resume with the saved handle
if (activeSessionHandle) {
  session = await liveModel.connect({
    handle: activeSessionHandle
  });
}

// Option 2: Resume the session directly on an existing session object
if (activeSessionHandle) {
  await session.resumeSession({
    handle: activeSessionHandle
  });
}

Dart

// Local variable to save the active session handle
String? _activeSessionHandle;

// Initialize the session. Passing an empty config requests the server to send SessionResumptionUpdate
var _session = await _liveModel.connect(
  sessionResumption: SessionResumptionConfig(),
);

// Start receiving responses
await for (final message in _session.receive()) {
  // Process other received response types...

  // Check for new session handles inside your message handling loop
  if (message is SessionResumptionUpdate &&
      message.resumable != null &&
      message.resumable!) {
    _activeSessionHandle = message.newHandle;
    log('SessionResumptionUpdate: handle ${message.newHandle}');
  }
}

// The following are alternative options to resume a session. Choose only one.

// Option 1: Create and connect a session to resume with the saved handle
if (_activeSessionHandle != null) {
  _session = await _liveModel.connect(
    sessionResumption: SessionResumptionConfig.resume(_activeSessionHandle!),
  );
}

// Option 2: Alternatively, resume the session directly on an existing session object
if (_activeSessionHandle != null) {
  await _session.resumeSession(
    sessionResumption: SessionResumptionConfig.resume(_activeSessionHandle!),
  );
}

Unity

// Local variable to save the active session handle
string activeSessionHandle = null;

// Initialize the session. Passing an empty config requests the server to send SessionResumptionUpdate
var session = await liveModel.ConnectAsync(
    sessionResumption: new SessionResumptionConfig()
);

// Start receiving responses
await foreach (var response in session.ReceiveAsync())
{
  // Process other received response types...

  // Check for new session handles inside your message handling loop
  if (response.Message is LiveSessionResumptionUpdate updateMessage)
  {
    if (updateMessage.Resumable == true && !string.IsNullOrEmpty(updateMessage.NewHandle))
    {
      activeSessionHandle = updateMessage.NewHandle;
      Debug.Log($"SessionResumptionUpdate: handle {activeSessionHandle}");
    }
  }
}

// The following are alternative options to resume a session. Choose only one.

// Option 1: Create and connect a session to resume with the saved handle
if (!string.IsNullOrEmpty(activeSessionHandle)) {
  session = await liveModel.ConnectAsync(
      sessionResumption: new SessionResumptionConfig(activeSessionHandle)
  );
}

// Option 2: Resume the session directly on an existing session object
if (!string.IsNullOrEmpty(activeSessionHandle)) {
  await session.ResumeSessionAsync(
      sessionResumption: new SessionResumptionConfig(activeSessionHandle)
  );
}