Extract response format
Field-by-field reference for the Extract v2 job response — the job envelope, the extract_result object shaped by your data schema, extract_metadata citations and confidence scores, extraction-target shapes, saved configurations, and usage credits.
An Extract job returns one object, the extract job, from POST /api/v2/extract,
GET /api/v2/extract/{job_id}, and the cancel endpoint. The SDKs return the same object from
client.extract.create(), client.extract.get(), and client.extract.wait_for_completion().
This page lists every field in that object and what each one means. For setting up the job, see
Configuring Extract; for the request side of the
reference, see the Extract API reference.
Job envelope
Section titled “Job envelope”| Field | Type | Description |
|---|---|---|
id | string | Job identifier, prefixed ext-. Pass it to GET /api/v2/extract/{job_id}. |
status | string | Job state. See Job status. |
file_input | string | The file ID (dfl-...) or parse job ID (pjb-...) the job ran on. |
project_id | string | Project the job belongs to. |
created_at, updated_at | datetime | Creation and last-update timestamps. |
error_message | string or null | Error details when status is FAILED. |
extract_result | object, array, or null | The extracted data, shaped by your data_schema. See extract_result. |
extract_metadata | object or null | Citations, confidence scores, and parse details. Requires expand=extract_metadata. See extract_metadata. |
configuration | object or null | The configuration the job ran with. Requires expand=configuration. See Saved configurations. |
configuration_id | string or null | Saved configuration ID used for the job, if any. |
usage | object or null | Credits billed against the job. Requires expand=usage. See Usage and credits. |
metadata | object or null | Job-level metadata. None of the documented expand values return it, so expect null. |
Expanding the response
Section titled “Expanding the response”GET /api/v2/extract/{job_id} returns extract_result by default and omits the heavier fields.
Add one or more expand values to include them. The create and cancel responses take no expand
and always include configuration.
expand value | Adds |
|---|---|
extract_metadata | extract_metadata — per-field citations and confidence scores |
configuration | configuration — the resolved configuration the job ran with |
usage | usage — credits billed against the job |
The list endpoint, GET /api/v2/extract, accepts expand=configuration and
expand=extract_metadata and returns items, a next_page_token for the next page, and an
optional total_size.
Job status
Section titled “Job status”status | Meaning |
|---|---|
PENDING | Queued, not yet started. |
RUNNING | Actively processing. |
COMPLETED | Finished; extract_result is populated. |
FAILED | Terminated with an error; read error_message. |
CANCELLED | Cancelled by the user. |
Poll until the status is one of the three terminal values, or let
client.extract.wait_for_completion(job.id) do it for you.
extract_result
Section titled “extract_result”extract_result is the data itself, and its shape is set by two things: your data_schema
decides the keys and value types, and extraction_target decides whether you get one object or an
array of them.
extraction_target | extract_result shape |
|---|---|
per_doc (default) | A single JSON object matching your schema. |
per_page | An array of objects, one per page, each matching your schema. |
per_table_row | An array of objects, one per detected entity (table row, list item, section), each matching your schema. |
A per_doc result for the schema in the
citations example:
{ "company_name": "NVIDIA Corporation", "filing_type": "10 K", "filing_date": "February 26, 2025", "fiscal_year": 2025, "unit": "millions", "revenue": 130497}Repeating entities
Section titled “Repeating entities”With per_table_row the schema describes a single entity and the response is an array with one
element per entity, so len(job.extract_result) is the entity count. The
repeating entities example returns
this shape for a hospital directory:
[ {"county": "Alameda", "hospital_name": "Alameda Hospital", "plan_names": ["Trio HMO", "SaveNet", "Access+ HMO"]}, {"county": "Alameda", "hospital_name": "Eden Medical Center", "plan_names": ["Trio HMO", "Access+ HMO", "PPO"]}]Arrays declared inside your schema (for example skills: list[str] or items: list[LineItem])
appear as ordinary JSON arrays inside each result object, whatever the extraction target.
Missing values
Section titled “Missing values”A field the document has no value for is handled according to your schema. An optional field
(one not listed in required) is left out of the result. A required field, or one declared
nullable, is returned as null. See
Required and optional fields.
extract_metadata
Section titled “extract_metadata”extract_metadata is returned only with expand=extract_metadata. Its field_metadata entries
carry citations and confidence scores only when the matching extension is turned on in the
configuration. It has three fields:
| Field | Type | Description |
|---|---|---|
field_metadata | object or null | Per-field citations, confidence scores, and reasoning. See below. |
parse_job_id | string or null | The Parse job that produced the document text the extraction read. |
parse_tier | string or null | Parse tier used for that Parse job. |
Turbo jobs produce no parse output; see
Tiers. Jobs run with spreadsheet_mode
produce neither citations nor confidence scores.
field_metadata
Section titled “field_metadata”field_metadata mirrors extract_result, and which of its three keys is populated follows the
extraction target:
| Key | Populated when | Shape |
|---|---|---|
document_metadata | extraction_target is per_doc | One object keyed by field name. |
page_metadata | extraction_target is per_page | An array of per-field metadata objects, one per page. |
row_metadata | extraction_target is per_table_row | An array of per-field metadata objects, one per row. |
Inside each of those, the tree follows your schema: a scalar field maps to a metadata entry; an
array field maps to a list where each element holds the entries for that element’s sub-fields,
indexed by array position; a nested object holds entries for its sub-fields recursively. The
metadata for items[0].amount therefore lives at document_metadata.items[0].amount, not at
document_metadata.items[0].
A metadata entry carries up to four keys, depending on which extensions are on:
| Key | Enabled by | Description |
|---|---|---|
citation | cite_sources: true | Array of source locations for the value. See Citations. |
confidence | confidence_scores: true | Combined confidence, 0 to 1. The value to threshold on. |
parsing_confidence | confidence_scores: true | How well the relevant context was parsed from the source document. |
extraction_confidence | confidence_scores: true | How well the extracted value matches the schema field. |
The example from the API reference, with both extensions on, for a schema with vendor, total,
and an items array:
"extract_metadata": { "field_metadata": { "document_metadata": { "vendor": { "citation": [ { "page": 1, "matching_text": "Noisebridge", "bounding_boxes": [{"x": 72.0, "y": 96.0, "w": 120.0, "h": 14.0}] } ], "confidence": 1.0, "extraction_confidence": 1.0, "parsing_confidence": 1.0 }, "total": { "citation": [{"page": 1, "matching_text": "$10.00"}], "confidence": 1.0 }, "items": [ { "amount": {"citation": [{"page": 1, "matching_text": "$10.00"}], "confidence": 1.0}, "description": {"citation": [{"page": 1, "matching_text": "$10/month"}], "confidence": 0.998} } ] } }}Citations
Section titled “Citations”Set cite_sources: true in the configuration. Every leaf field then carries a citation array,
with one element per place the value was found, so a field cited across several pages has several
elements. Each element has:
| Key | Type | Description |
|---|---|---|
page | integer | 1-based page number where the text was found. |
matching_text | string | The verbatim source text the value was extracted from. |
bounding_boxes | array of {x, y, w, h} | Location of the cited text on the page. |
page_dimensions | {width, height} | Page size, for scaling the bounding boxes when you render them. |
On the turbo tier citations are text-level only: page and matching_text are returned
without bounding boxes.
Confidence scores
Section titled “Confidence scores”Set confidence_scores: true in the configuration. confidence combines parsing_confidence
and extraction_confidence and is the score to route on; the other two explain where a low score
came from. Calibration differs by tier, and long free-text fields score lower than short factual
ones; see Confidence scores for how to
pick a threshold.
Reasoning
Section titled “Reasoning”Per-field reasoning strings appear in field_metadata for Extract versions through
2026-03-31 and are not returned by newer versions; see
Reasoning metadata.
Saved configurations
Section titled “Saved configurations”A job created with configuration_id runs with the saved configuration’s parameters, so the
extensions, extraction target, and tier in that saved configuration decide what the response
contains. The response records which configuration was used:
configuration_idechoes the saved configuration ID. It is null for jobs created with an inlineconfiguration.configuration(withexpand=configuration) is the full resolved configuration the job ran with:data_schema,tier,version,extraction_target,cite_sources,confidence_scores, and the other options from Configuring Extract.configuration.versionis the concrete release the job ran, reported as a release name such as2.5and fixed at job creation. It is neverlatest, even when the saved configuration storeslatestor a date.
See Using saved configurations for creating and reusing them.
Usage and credits
Section titled “Usage and credits”usage, returned with expand=usage, describes what the job consumed:
| Key | Description |
|---|---|
credits | Total credits billed against the job: the sum of the two components below, counting only those already recorded. |
extract_credits | Credits billed for the extraction itself. |
parse_credits | Credits billed against the Parse job that Extract created for this job. null when file_input was a Parse job you created yourself, because those credits belong to that Parse job. |
usage is null until the job is COMPLETED, and each value inside it reads null until
billing has recorded it, which can trail job completion, so poll again rather than reading null
as zero. One Parse job can back several Extract jobs, and each of them reports that same
parse_credits; to total across jobs, sum extract_credits and count the parse once. See
Check the credits a job billed.
Fetch the full response
Section titled “Fetch the full response”import osfrom llama_cloud import LlamaCloud
client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])
DATA_SCHEMA = { "type": "object", "properties": { "company_name": {"type": "string", "description": "Name of the company"}, "revenue": {"type": "number", "description": "Annual revenue in USD"}, },}
file_obj = client.files.create(file="path/to/document.pdf", purpose="extract")
job = client.extract.create( file_input=file_obj.id, configuration={ "data_schema": DATA_SCHEMA, "tier": "agentic", "cite_sources": True, "confidence_scores": True, },)job = client.extract.wait_for_completion(job.id)
# Metadata, configuration, and usage are omitted unless expandedjob = client.extract.get(job.id, expand=["extract_metadata", "configuration", "usage"])
print(job.status)print(job.extract_result)print(job.extract_metadata.field_metadata.document_metadata)print(job.usage)