Turn any PDF into structured data with a schema you define.
Pull structured data out of PDFs using Google Gemini, with a schema you control.
Give it a document and a field list, and it returns clean JSON. No per-document
templates, no coordinate mapping, no retraining — the schema is the configuration,
so the same toolkit handles invoices, statements, forms, reports and contracts
without changes.
null ratherextract_pdf_fields with a storage path and your field list.The API key is read from the secret store at call time and is never returned in a
response or written to the extraction record.
{
"fields": [
{"name": "invoice_number", "type": "string", "description": "Invoice or reference number"},
{"name": "issue_date", "type": "date", "description": "Date of issue, ISO 8601"},
{"name": "total_amount", "type": "number", "description": "Total including tax"},
{"name": "currency", "type": "string", "description": "ISO 4217 code"},
{"name": "is_paid", "type": "boolean", "description": "Whether marked paid"}
]
}
Supported types are string, number, integer, , and .
booleandatearrayOne call per document. Large PDFs take longer and cost more — extraction runs with a
240-second timeout, and documents over roughly 50 pages are better split first.
Adding this toolkit deploys the following callable cloud functions onto your app:
classify_pdf()extract_pdf_fields()You'll be prompted to provide values for these secret types during install: