使用 Gemini API 分析文件 (例如 PDF)

您可以要求 Gemini 模型分析您提供的文件檔案 (例如 PDF 和純文字檔案),這些檔案可以內嵌 (採用 base64 編碼) 或透過網址提供。使用 Firebase AI Logic 時,您可以直接從應用程式提出這項要求。

這項功能可協助您執行下列操作:

  • 分析文件中的圖表和表格
  • 以結構化輸出格式擷取資訊
  • 回答文件中的圖片和文字內容相關問題
  • 生成文件摘要
  • 轉錄文件內容 (例如轉成 HTML),保留版面配置和格式,供下游應用程式使用 (例如 RAG 管道)

本指南說明如何從文件輸入內容 (例如 PDF) 生成文字,但您也可以從文件輸入內容生成圖片。

跳至程式碼範例 跳至串流回應的程式碼


如要瞭解處理文件 (例如 PDF) 的其他選項,請參閱其他指南
產生結構化輸出內容 多輪對話

事前準備

按一下 Gemini API 供應商,即可在這個頁面查看供應商專屬內容和程式碼。

如果尚未完成,請參閱入門指南,瞭解如何設定 Firebase 專案、將應用程式連結至 Firebase、新增 SDK、為所選Gemini API供應商初始化後端服務,以及建立 GenerativeModel 執行個體。

如要測試及反覆調整提示,建議使用 Google AI Studio。

支援這項功能的機型

本指南說明如何從文件輸入內容 (例如 PDF) 生成文字,適用於下列 Gemini 模型:

  • gemini-3.1-pro-preview
  • gemini-3.8-flash (以及舊版 gemini-3.7-flash、gemini-3.6-flash 和 gemini-3.5-flash)
  • gemini-3.5-flash-lite (和舊版 gemini-3.1-flash-lite)

一般用途 Gemini 2.5 模型支援這項功能,但都已淘汰。

從 PDF 檔案 (採用 Base64 編碼) 生成文字

嘗試這個範例前,請先完成本指南的「事前準備」一節,設定專案和應用程式。
在該節中,您也會點選所選Gemini API供應商的按鈕,以便在本頁面查看供應商專屬內容。

您可以透過文字和 PDF 提示 Gemini 模型生成文字,方法是提供每個輸入檔案的 mimeType 和檔案本身。請參閱本頁面稍後的輸入檔案規定和建議。

Swift

您可以呼叫 generateContent(),從文字和 PDF 的多模態輸入內容生成文字。


import FirebaseAILogic

// Initialize the Gemini Developer API backend service.
let ai = FirebaseAI.firebaseAI(backend: .googleAI())

// Create a `GenerativeModel` instance with a model that supports your use case.
let model = ai.generativeModel(modelName: "gemini-3.8-flash")


// Provide the PDF as `Data` with the appropriate MIME type
let pdf = try InlineDataPart(data: Data(contentsOf: pdfURL), mimeType: "application/pdf")

// Provide a text prompt to include with the PDF file
let prompt = "Summarize the important results in this report."

// To generate text output, call `generateContent` with the PDF file and text prompt
let response = try await model.generateContent(pdf, prompt)

// Print the generated text, handling the case where it might be nil
print(response.text ?? "No text in response.")

Kotlin

您可以呼叫 generateContent(),從文字和 PDF 的多模態輸入內容生成文字。

如果是 Kotlin,這個 SDK 中的方法是暫停函式,需要從 Coroutine 範圍呼叫。

// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
val model = Firebase.ai(backend = GenerativeBackend.googleAI())
                        .generativeModel("gemini-3.8-flash")


val contentResolver = applicationContext.contentResolver

// Provide the URI for the PDF file you want to send to the model
val inputStream = contentResolver.openInputStream(pdfUri)

if (inputStream != null) {  // Check if the PDF file loaded successfully
    inputStream.use { stream ->
        // Provide a prompt that includes the PDF file specified above and text
        val prompt = content {
            inlineData(
                bytes = stream.readBytes(),
                mimeType = "application/pdf" // Specify the appropriate PDF file MIME type
            )
            text("Summarize the important results in this report.")
        }

        // To generate text output, call `generateContent` with the prompt
        val response = model.generateContent(prompt)

        // Log the generated text, handling the case where it might be null
        Log.d(TAG, response.text ?: "")
    }
} else {
    Log.e(TAG, "Error getting input stream for file.")
    // Handle the error appropriately
}

Java

您可以呼叫 generateContent(),從文字和 PDF 的多模態輸入內容生成文字。

如果是 Java,這個 SDK 中的方法會傳回 ListenableFuture。

// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI())
        .generativeModel("gemini-3.8-flash");

// Use the GenerativeModelFutures Java compatibility layer which offers
// support for ListenableFuture and Publisher APIs
GenerativeModelFutures model = GenerativeModelFutures.from(ai);


ContentResolver resolver = getApplicationContext().getContentResolver();

// Provide the URI for the PDF file you want to send to the model
try (InputStream stream = resolver.openInputStream(pdfUri)) {
    if (stream != null) {
        byte[] audioBytes = stream.readAllBytes();
        stream.close();

        // Provide a prompt that includes the PDF file specified above and text
        Content prompt = new Content.Builder()
              .addInlineData(audioBytes, "application/pdf")  // Specify the appropriate PDF file MIME type
              .addText("Summarize the important results in this report.")
              .build();

        // To generate text output, call `generateContent` with the prompt
        ListenableFuture<GenerateContentResponse> response = model.generateContent(prompt);
        Futures.addCallback(response, new FutureCallback<GenerateContentResponse>() {
            @Override
            public void onSuccess(GenerateContentResponse result) {
                String text = result.getText();
                Log.d(TAG, (text == null) ? "" : text);
            }
            @Override
            public void onFailure(Throwable t) {
                Log.e(TAG, "Failed to generate a response", t);
            }
        }, executor);
    } else {
        Log.e(TAG, "Error getting input stream for file.");
        // Handle the error appropriately
    }
} catch (IOException e) {
    Log.e(TAG, "Failed to read the pdf file", e);
} catch (URISyntaxException e) {
    Log.e(TAG, "Invalid pdf file", e);
}

Web

您可以呼叫 generateContent(),從文字和 PDF 的多模態輸入內容生成文字。


import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend } from "firebase/ai";

// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
  // ...
};

// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);

// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });

// Create a `GenerativeModel` instance with a model that supports your use case.
const model = getGenerativeModel(ai, { model: "gemini-3.8-flash" });


// Converts a File object to a Part object.
async function fileToGenerativePart(file) {
  const base64EncodedDataPromise = new Promise((resolve) => {
    const reader = new FileReader();
    reader.onloadend = () => resolve(reader.result.split(','));
    reader.readAsDataURL(file);
  });
  return {
    inlineData: { data: await base64EncodedDataPromise, mimeType: file.type },
  };
}

async function run() {
  // Provide a text prompt to include with the PDF file
  const prompt = "Summarize the important results in this report.";

  // Prepare PDF file for input
  const fileInputEl = document.querySelector("input[type=file]");
  const pdfPart = await fileToGenerativePart(fileInputEl.files);

  // To generate text output, call `generateContent` with the text and PDF file
  const result = await model.generateContent([prompt, pdfPart]);

  // Log the generated text, handling the case where it might be undefined
  console.log(result.response.text() ?? "No text in response.");
}

run();

Dart

您可以呼叫 generateContent(),從文字和 PDF 的多模態輸入內容生成文字。


import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';

// Initialize FirebaseApp
await Firebase.initializeApp(
  options: DefaultFirebaseOptions.currentPlatform,
);

// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
final model =
      FirebaseAI.googleAI().generativeModel(model: 'gemini-3.8-flash');


// Provide a text prompt to include with the PDF file
final prompt = TextPart("Summarize the important results in this report.");

// Prepare the PDF file for input
final doc = await File('document0.pdf').readAsBytes();

// Provide the PDF file as `Data` with the appropriate PDF file MIME type
final docPart = InlineDataPart('application/pdf', doc);

// To generate text output, call `generateContent` with the text and PDF file
final response = await model.generateContent([
  Content.multi([prompt,docPart])
]);

// Print the generated text
print(response.text);

Unity

您可以呼叫 GenerateContentAsync(),從文字和 PDF 的多模態輸入內容生成文字。


using Firebase;
using Firebase.AI;

// Initialize the Gemini Developer API backend service.
var ai = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI());

// Create a `GenerativeModel` instance with a model that supports your use case.
var model = ai.GetGenerativeModel(modelName: "gemini-3.8-flash");


// Provide a text prompt to include with the PDF file
var prompt = ModelContent.Text("Summarize the important results in this report.");

// Provide the PDF file as `data` with the appropriate PDF file MIME type
var doc = ModelContent.InlineData("application/pdf",
      System.IO.File.ReadAllBytes(System.IO.Path.Combine(
        UnityEngine.Application.streamingAssetsPath, "document0.pdf")));

// To generate text output, call `GenerateContentAsync` with the text and PDF file
var response = await model.GenerateContentAsync(new [] { prompt, doc });

// Print the generated text
UnityEngine.Debug.Log(response.Text ?? "No text in response.");

瞭解如何選擇適合應用程式和用途的模型, 。

逐句顯示回覆

嘗試這個範例前,請先完成本指南的「事前準備」一節,設定專案和應用程式。
在該節中,您也會點選所選Gemini API供應商的按鈕,以便在本頁面查看供應商專屬內容。

您不必等待模型生成完整結果,而是使用串流處理部分結果,即可加快互動速度。如要串流回覆,請呼叫 generateContentStream。



輸入文件的規定和建議

請注意,以內嵌資料形式提供的檔案在傳輸過程中會編碼為 base64,這會增加要求的大小。如果要求過大,就會收到 HTTP 413 錯誤。

如要進一步瞭解下列事項,請參閱「支援的輸入檔案和規定」頁面:

支援的文件 MIME 類型

Gemini 多模態模型支援下列文件 MIME 類型:

  • PDF - application/pdf
  • 傳送訊息到 text/plain

每項要求的限制

PDF 會視為圖片,因此一個 PDF 頁面相當於一張圖片。提示中允許的頁面數量,取決於 Gemini 多模態模型支援的圖片數量。

  • 每項要求的檔案數量上限:3,000 個檔案
  • 每個檔案的頁數上限:1,000 頁
  • 每個檔案的大小上限:50 MB



你還可以做些什麼?

試試其他功能

瞭解如何控管內容生成功能

您也可以使用 Google AI Studio 測試提示和模型設定,甚至取得生成的程式碼片段。

進一步瞭解支援的機型

瞭解各種用途適用的模型,以及這些模型的配額和價格。


提供有關 Firebase AI Logic 的使用體驗意見回饋