حتی با پیادهسازی پایه برای Live API، میتوانید تعاملهای جذاب و قدرتمندی برای کاربران خود بسازید. بااستفاده از گزینههای پیکربندی زیر میتوانید بهصورت اختیاری تجربه را حتی بیشتر سفارشیسازی کنید:
زبان و صدای پاسخ
میتوانید مدل را وادار کنید با صدای خاصی پاسخ دهد و بر مدل تأثیر بگذارید تا به زبانهای مختلف پاسخ دهد.
صدای پاسخ را مشخص کنید
|
روی ارائهدهنده Gemini API خود کلیک کنید تا محتوا و کد مخصوص ارائهدهنده را در این صفحه مشاهده کنید. |
Live API از ۳۰ صدای HD مصنوعی مختلف پشتیبانی میکند که هرکدام ویژگیهای متمایز دارند. با ازهم بازکردن بخش زیر، میتوانید فهرست گزینههای صدای پاسخ را ببینید و صدای نمایشی هر صدا را بشنوید
اگر صدای پاسخ را مشخص نکنید، صدای پیشفرض Puck است.
برای مشخص کردن صدای پاسخ، نام صدا را در speechConfig شیء
بهعنوان بخشی از
پیکربندی مدل تنظیم کنید.
Swift
// ...
let liveModel = FirebaseAI.firebaseAI(backend: .googleAI()).liveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to use a specific voice for its audio response.
generationConfig: LiveGenerationConfig(
responseModalities: [.audio],
speech: SpeechConfig(voiceName: "VOICE_NAME")
)
)
// ...
Kotlin
// ...
val model = Firebase.ai(backend = GenerativeBackend.googleAI()).liveModel(
modelName = "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to use a specific voice for its audio response.
generationConfig = liveGenerationConfig {
responseModality = ResponseModality.AUDIO
speechConfig = SpeechConfig(voice = Voice("VOICE_NAME"))
}
)
// ...
Java
// ...
LiveGenerativeModel lm = FirebaseAI.getInstance(GenerativeBackend.googleAI()).liveModel(
"GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to use a specific voice for its audio response.
new LiveGenerationConfig.Builder()
.setResponseModality(ResponseModality.AUDIO)
.setSpeechConfig(new SpeechConfig(new Voice("VOICE_NAME")))
.build()
);
// ...
Web
// ...
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
const liveModel = getLiveGenerativeModel(ai, {
model: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to use a specific voice for its audio response.
generationConfig: {
responseModalities: [ResponseModality.AUDIO],
speechConfig: {
voiceConfig: {
prebuiltVoiceConfig: { voiceName: "VOICE_NAME" },
},
},
},
});
// ...
Dart
// ...
final _liveModel = FirebaseAI.googleAI().liveGenerativeModel(
model: 'GEMINI_LIVE_API_MODEL_NAME',
// Configure the model to use a specific voice for its audio response.
liveGenerationConfig: LiveGenerationConfig(
responseModalities: [ResponseModalities.audio],
speechConfig: SpeechConfig(voiceName: 'VOICE_NAME'),
),
);
// ...
Unity
// ...
var liveModel = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetLiveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to use a specific voice for its audio response
liveGenerationConfig: new LiveGenerationConfig(
responseModalities: new[] { ResponseModality.Audio },
speechConfig: SpeechConfig.UsePrebuiltVoice("VOICE_NAME")
)
);
// ...
تأثیرگذاری بر زبان پاسخ
مدلهای Live API بهطور خودکار زبان مناسب را برای پاسخهایشان انتخاب میکنند.
اگر میخواهید مدل به زبانی غیراز انگلیسی یا همیشه به زبانی خاص پاسخ دهد، میتوانید بااستفاده از دستورالعملهای سیستم مانند این مثالها، بر پاسخهای مدل تأثیر بگذارید:
به مدل تأکید کنید که زبان غیرانگلیسی ممکن است مناسب باشد
Listen to the speaker carefully. If you detect a non-English language, respond in the language you hear from the speaker. You must respond unmistakably in the speaker's language.به مدل بگویید همیشه به زبان خاصی پاسخ دهد
RESPOND IN LANGUAGE. YOU MUST RESPOND UNMISTAKABLY IN LANGUAGE.
ترانویسی برای ورودی و خروجی صوتی
|
روی ارائهدهنده Gemini API خود کلیک کنید تا محتوا و کد مخصوص ارائهدهنده را در این صفحه مشاهده کنید. |
بهعنوان بخشی از پاسخ مدل، میتوانید ترانویسیهای ورودی صوتی و پاسخ صوتی مدل را دریافت کنید. این پیکربندی را بهعنوان بخشی از پیکربندی مدل تنظیم میکنید.
برای ترانویسی ورودی صوتی،
inputAudioTranscriptionرا اضافه کنید.برای ترانویسی پاسخ صوتی مدل،
outputAudioTranscriptionرا اضافه کنید.
به موارد زیر توجه کنید:
میتوانید مدل را پیکربندی کنید تا ترانویسیهای ورودی و خروجی را برگرداند (همانطور که در مثال زیر نشان داده شده است)، یا میتوانید آن را پیکربندی کنید تا فقط یکی از آنها را برگرداند.
ترانویسیها همراه با صدا جاریسازی میشوند، بنابراین بهتر است آنها را مانند بخشهای نوشتاری در هر نوبت جمعآوری کنید.
زبان ترانویسی از ورودی صوتی و پاسخ صوتی مدل استنباط میشود.
Swift
// ...
let liveModel = FirebaseAI.firebaseAI(backend: .googleAI()).liveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to return transcriptions of the audio input and output.
generationConfig: LiveGenerationConfig(
responseModalities: [.audio],
inputAudioTranscription: AudioTranscriptionConfig(),
outputAudioTranscription: AudioTranscriptionConfig()
)
)
var inputTranscript: String = ""
var outputTranscript: String = ""
do {
let session = try await liveModel.connect()
for try await response in session.responses {
if case let .content(content) = response.payload {
if let inputText = content.inputAudioTranscription?.text {
// Handle transcription text of the audio input.
inputTranscript += inputText
}
if let outputText = content.outputAudioTranscription?.text {
// Handle transcription text of the audio output.
outputTranscript += outputText
}
if content.isTurnComplete {
// Log the transcripts after the current turn is complete.
print("Input audio: \(inputTranscript)")
print("Output audio: \(outputTranscript)")
// Reset the transcripts for the next turn.
inputTranscript = ""
outputTranscript = ""
}
}
}
} catch {
// Handle error
}
// ...
Kotlin
// ...
val liveModel = Firebase.ai(backend = GenerativeBackend.googleAI()).liveModel(
modelName = "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to return transcriptions of the audio input and output.
generationConfig = liveGenerationConfig {
responseModality = ResponseModality.AUDIO
inputAudioTranscription = AudioTranscriptionConfig()
outputAudioTranscription = AudioTranscriptionConfig()
}
)
val liveSession = liveModel.connect()
fun handleTranscription(input: Transcription?, output: Transcription?) {
input?.text?.let { text ->
// Handle transcription text of the audio input.
println("Input Transcription: $text")
}
output?.text?.let { text ->
// Handle transcription text of the audio output.
println("Output Transcription: $text")
}
}
liveSession.startAudioConversation(null, ::handleTranscription)
// ...
Java
// ...
ExecutorService executor = Executors.newFixedThreadPool(1);
LiveGenerativeModel lm = FirebaseAI.getInstance(GenerativeBackend.googleAI()).liveModel(
"GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to return transcriptions of the audio input and output.
new LiveGenerationConfig.Builder()
.setResponseModality(ResponseModality.AUDIO)
.setInputAudioTranscription(new AudioTranscriptionConfig())
.setOutputAudioTranscription(new AudioTranscriptionConfig())
.build()
);
LiveModelFutures liveModel = LiveModelFutures.from(lm);
ListenableFuture<LiveSessionFutures> sessionFuture = liveModel.connect();
Futures.addCallback(sessionFuture, new FutureCallback<LiveSessionFutures>() {
@Override
public void onSuccess(LiveSessionFutures ses) {
LiveSessionFutures session = ses;
session.startAudioConversation((Transcription input, Transcription output) -> {
if (input != null) {
// Handle transcription text of the audio input.
System.out.println("Input Transcription: " + input.getText());
}
if (output != null) {
// Handle transcription text of the audio output.
System.out.println("Output Transcription: " + output.getText());
}
return null;
});
}
@Override
public void onFailure(Throwable t) {
// Handle exceptions
t.printStackTrace();
}
}, executor);
// ...
Web
// ...
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
const liveModel = getLiveGenerativeModel(ai, {
model: 'GEMINI_LIVE_API_MODEL_NAME',
// Configure the model to return transcriptions of the audio input and output.
generationConfig: {
responseModalities: [ResponseModality.AUDIO],
inputAudioTranscription: {},
outputAudioTranscription: {},
},
});
const liveSession = await liveModel.connect();
liveSession.sendAudioRealtime({ data, mimeType: "audio/pcm" });
const messages = liveSession.receive();
for await (const message of messages) {
switch (message.type) {
case 'serverContent':
if (message.inputTranscription) {
// Handle transcription text of the audio input.
console.log(`Input transcription: ${message.inputTranscription.text}`);
}
if (message.outputTranscription) {
// Handle transcription text of the audio output.
console.log(`Output transcription: ${message.outputTranscription.text}`);
} else {
// Handle other message types (modelTurn, turnComplete, interruption).
}
default:
// Handle other message types (toolCall, toolCallCancellation).
}
}
// ...
Dart
// ...
final _liveModel = FirebaseAI.googleAI().liveGenerativeModel(
model: 'GEMINI_LIVE_API_MODEL_NAME',
// Configure the model to return transcriptions of the audio input and output.
liveGenerationConfig: LiveGenerationConfig(
responseModalities: [ResponseModalities.audio],
inputAudioTranscription: AudioTranscriptionConfig(),
outputAudioTranscription: AudioTranscriptionConfig(),
),
);
final LiveSession _session = _liveModel.connect();
await for (final response in _session.receive()) {
LiveServerContent message = response.message;
if (message.inputTranscription?.text case final inputText?) {
// Handle transcription text of the audio input.
print('Input: $inputText');
}
if (message.outputTranscription?.text case final outputText?) {
// Handle transcription text of the audio output.
print('Output: $outputText');
}
}
// ...
Unity
// ...
var liveModel = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetLiveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to return transcriptions of the audio input and output
liveGenerationConfig: new LiveGenerationConfig(
responseModalities: new[] { ResponseModality.Audio },
inputAudioTranscription: new AudioTranscriptionConfig(),
outputAudioTranscription: new AudioTranscriptionConfig()
)
);
try
{
var session = await liveModel.ConnectAsync();
var stream = session.ReceiveAsync();
await foreach (var response in stream) {
if (response.Message is LiveSessionContent sessionContent) {
if (!string.IsNullOrEmpty(sessionContent.InputTranscription?.Text)) {
// handle transcription text of input audio
}
if (!string.IsNullOrEmpty(sessionContent.OutputTranscription?.Text)) {
// handle transcription text of output audio
}
}
}
}
catch (Exception e)
{
// Handle error
}
// ...
تشخیص فعالیت گفتاری (VAD)
این مدل بهطور خودکار تشخیص فعالیت گفتاری (VAD) را روی جاریسازی ورودی صوتی پیوسته انجام میدهد. «تشخیص فعالیت صوتی» بهطور پیشفرض فعال است.
مدیریت جلسه
درباره موضوعات مربوط به جلسات زیر بیشتر بدانید:
قابلیتهای پیشرفته، ازجمله:
محدودیتهای مربوط به جلسه، ازجمله محدودیتهای اتصال و طول جلسه، محدودیتهای پنجره بافت جلسه، و محدودیتهای نرخ.
گزینههایی برای مدیریت محدودیتهای جلسه، ازجمله: