Skip to content
Guide
Extract
Guides

Extract response format

Field-by-field reference for the Extract v2 job response — the job envelope, the extract_result object shaped by your data schema, extract_metadata citations and confidence scores, extraction-target shapes, saved configurations, and usage credits.

An Extract job returns one object, the extract job, from POST /api/v2/extract, GET /api/v2/extract/{job_id}, and the cancel endpoint. The SDKs return the same object from client.extract.create(), client.extract.get(), and client.extract.wait_for_completion(). This page lists every field in that object and what each one means. For setting up the job, see Configuring Extract; for the request side of the reference, see the Extract API reference.

FieldTypeDescription
idstringJob identifier, prefixed ext-. Pass it to GET /api/v2/extract/{job_id}.
statusstringJob state. See Job status.
file_inputstringThe file ID (dfl-...) or parse job ID (pjb-...) the job ran on.
project_idstringProject the job belongs to.
created_at, updated_atdatetimeCreation and last-update timestamps.
error_messagestring or nullError details when status is FAILED.
extract_resultobject, array, or nullThe extracted data, shaped by your data_schema. See extract_result.
extract_metadataobject or nullCitations, confidence scores, and parse details. Requires expand=extract_metadata. See extract_metadata.
configurationobject or nullThe configuration the job ran with. Requires expand=configuration. See Saved configurations.
configuration_idstring or nullSaved configuration ID used for the job, if any.
usageobject or nullCredits billed against the job. Requires expand=usage. See Usage and credits.
metadataobject or nullJob-level metadata. None of the documented expand values return it, so expect null.

GET /api/v2/extract/{job_id} returns extract_result by default and omits the heavier fields. Add one or more expand values to include them. The create and cancel responses take no expand and always include configuration.

expand valueAdds
extract_metadataextract_metadata — per-field citations and confidence scores
configurationconfiguration — the resolved configuration the job ran with
usageusage — credits billed against the job

The list endpoint, GET /api/v2/extract, accepts expand=configuration and expand=extract_metadata and returns items, a next_page_token for the next page, and an optional total_size.

statusMeaning
PENDINGQueued, not yet started.
RUNNINGActively processing.
COMPLETEDFinished; extract_result is populated.
FAILEDTerminated with an error; read error_message.
CANCELLEDCancelled by the user.

Poll until the status is one of the three terminal values, or let client.extract.wait_for_completion(job.id) do it for you.

extract_result is the data itself, and its shape is set by two things: your data_schema decides the keys and value types, and extraction_target decides whether you get one object or an array of them.

extraction_targetextract_result shape
per_doc (default)A single JSON object matching your schema.
per_pageAn array of objects, one per page, each matching your schema.
per_table_rowAn array of objects, one per detected entity (table row, list item, section), each matching your schema.

A per_doc result for the schema in the citations example:

{
"company_name": "NVIDIA Corporation",
"filing_type": "10 K",
"filing_date": "February 26, 2025",
"fiscal_year": 2025,
"unit": "millions",
"revenue": 130497
}

With per_table_row the schema describes a single entity and the response is an array with one element per entity, so len(job.extract_result) is the entity count. The repeating entities example returns this shape for a hospital directory:

[
{"county": "Alameda", "hospital_name": "Alameda Hospital", "plan_names": ["Trio HMO", "SaveNet", "Access+ HMO"]},
{"county": "Alameda", "hospital_name": "Eden Medical Center", "plan_names": ["Trio HMO", "Access+ HMO", "PPO"]}
]

Arrays declared inside your schema (for example skills: list[str] or items: list[LineItem]) appear as ordinary JSON arrays inside each result object, whatever the extraction target.

A field the document has no value for is handled according to your schema. An optional field (one not listed in required) is left out of the result. A required field, or one declared nullable, is returned as null. See Required and optional fields.

extract_metadata is returned only with expand=extract_metadata. Its field_metadata entries carry citations and confidence scores only when the matching extension is turned on in the configuration. It has three fields:

FieldTypeDescription
field_metadataobject or nullPer-field citations, confidence scores, and reasoning. See below.
parse_job_idstring or nullThe Parse job that produced the document text the extraction read.
parse_tierstring or nullParse tier used for that Parse job.

Turbo jobs produce no parse output; see Tiers. Jobs run with spreadsheet_mode produce neither citations nor confidence scores.

field_metadata mirrors extract_result, and which of its three keys is populated follows the extraction target:

KeyPopulated whenShape
document_metadataextraction_target is per_docOne object keyed by field name.
page_metadataextraction_target is per_pageAn array of per-field metadata objects, one per page.
row_metadataextraction_target is per_table_rowAn array of per-field metadata objects, one per row.

Inside each of those, the tree follows your schema: a scalar field maps to a metadata entry; an array field maps to a list where each element holds the entries for that element’s sub-fields, indexed by array position; a nested object holds entries for its sub-fields recursively. The metadata for items[0].amount therefore lives at document_metadata.items[0].amount, not at document_metadata.items[0].

A metadata entry carries up to four keys, depending on which extensions are on:

KeyEnabled byDescription
citationcite_sources: trueArray of source locations for the value. See Citations.
confidenceconfidence_scores: trueCombined confidence, 0 to 1. The value to threshold on.
parsing_confidenceconfidence_scores: trueHow well the relevant context was parsed from the source document.
extraction_confidenceconfidence_scores: trueHow well the extracted value matches the schema field.

The example from the API reference, with both extensions on, for a schema with vendor, total, and an items array:

"extract_metadata": {
"field_metadata": {
"document_metadata": {
"vendor": {
"citation": [
{
"page": 1,
"matching_text": "Noisebridge",
"bounding_boxes": [{"x": 72.0, "y": 96.0, "w": 120.0, "h": 14.0}]
}
],
"confidence": 1.0,
"extraction_confidence": 1.0,
"parsing_confidence": 1.0
},
"total": {
"citation": [{"page": 1, "matching_text": "$10.00"}],
"confidence": 1.0
},
"items": [
{
"amount": {"citation": [{"page": 1, "matching_text": "$10.00"}], "confidence": 1.0},
"description": {"citation": [{"page": 1, "matching_text": "$10/month"}], "confidence": 0.998}
}
]
}
}
}

Set cite_sources: true in the configuration. Every leaf field then carries a citation array, with one element per place the value was found, so a field cited across several pages has several elements. Each element has:

KeyTypeDescription
pageinteger1-based page number where the text was found.
matching_textstringThe verbatim source text the value was extracted from.
bounding_boxesarray of {x, y, w, h}Location of the cited text on the page.
page_dimensions{width, height}Page size, for scaling the bounding boxes when you render them.

On the turbo tier citations are text-level only: page and matching_text are returned without bounding boxes.

Set confidence_scores: true in the configuration. confidence combines parsing_confidence and extraction_confidence and is the score to route on; the other two explain where a low score came from. Calibration differs by tier, and long free-text fields score lower than short factual ones; see Confidence scores for how to pick a threshold.

Per-field reasoning strings appear in field_metadata for Extract versions through 2026-03-31 and are not returned by newer versions; see Reasoning metadata.

A job created with configuration_id runs with the saved configuration’s parameters, so the extensions, extraction target, and tier in that saved configuration decide what the response contains. The response records which configuration was used:

  • configuration_id echoes the saved configuration ID. It is null for jobs created with an inline configuration.
  • configuration (with expand=configuration) is the full resolved configuration the job ran with: data_schema, tier, version, extraction_target, cite_sources, confidence_scores, and the other options from Configuring Extract.
  • configuration.version is the concrete release the job ran, reported as a release name such as 2.5 and fixed at job creation. It is never latest, even when the saved configuration stores latest or a date.

See Using saved configurations for creating and reusing them.

usage, returned with expand=usage, describes what the job consumed:

KeyDescription
creditsTotal credits billed against the job: the sum of the two components below, counting only those already recorded.
extract_creditsCredits billed for the extraction itself.
parse_creditsCredits billed against the Parse job that Extract created for this job. null when file_input was a Parse job you created yourself, because those credits belong to that Parse job.

usage is null until the job is COMPLETED, and each value inside it reads null until billing has recorded it, which can trail job completion, so poll again rather than reading null as zero. One Parse job can back several Extract jobs, and each of them reports that same parse_credits; to total across jobs, sum extract_credits and count the parse once. See Check the credits a job billed.

import os
from llama_cloud import LlamaCloud
client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])
DATA_SCHEMA = {
"type": "object",
"properties": {
"company_name": {"type": "string", "description": "Name of the company"},
"revenue": {"type": "number", "description": "Annual revenue in USD"},
},
}
file_obj = client.files.create(file="path/to/document.pdf", purpose="extract")
job = client.extract.create(
file_input=file_obj.id,
configuration={
"data_schema": DATA_SCHEMA,
"tier": "agentic",
"cite_sources": True,
"confidence_scores": True,
},
)
job = client.extract.wait_for_completion(job.id)
# Metadata, configuration, and usage are omitted unless expanded
job = client.extract.get(job.id, expand=["extract_metadata", "configuration", "usage"])
print(job.status)
print(job.extract_result)
print(job.extract_metadata.field_metadata.document_metadata)
print(job.usage)
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/