---
title: Extract response format | Developer Documentation
description: Field-by-field reference for the Extract v2 job response — the job envelope, the extract_result object shaped by your data schema, extract_metadata citations and confidence scores, extraction-target shapes, saved configurations, and usage credits.
---

An Extract job returns one object, the extract job, from `POST /api/v2/extract`, `GET /api/v2/extract/{job_id}`, and the cancel endpoint. The SDKs return the same object from `client.extract.create()`, `client.extract.get()`, and `client.extract.wait_for_completion()`. This page lists every field in that object and what each one means. For setting up the job, see [Configuring Extract](/llamaparse/extract/guides/configuring-extract/index.md); for the request side of the reference, see the [Extract API reference](https://developers.llamaindex.ai/reference/resources/extract/).

## Job envelope

| Field                      | Type                   | Description                                                                                                                      |
| -------------------------- | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `id`                       | string                 | Job identifier, prefixed `ext-`. Pass it to `GET /api/v2/extract/{job_id}`.                                                      |
| `status`                   | string                 | Job state. See [Job status](#job-status).                                                                                        |
| `file_input`               | string                 | The file ID (`dfl-...`) or parse job ID (`pjb-...`) the job ran on.                                                              |
| `project_id`               | string                 | Project the job belongs to.                                                                                                      |
| `created_at`, `updated_at` | datetime               | Creation and last-update timestamps.                                                                                             |
| `error_message`            | string or null         | Error details when `status` is `FAILED`.                                                                                         |
| `extract_result`           | object, array, or null | The extracted data, shaped by your `data_schema`. See [extract\_result](#extract_result).                                        |
| `extract_metadata`         | object or null         | Citations, confidence scores, and parse details. Requires `expand=extract_metadata`. See [extract\_metadata](#extract_metadata). |
| `configuration`            | object or null         | The configuration the job ran with. Requires `expand=configuration`. See [Saved configurations](#saved-configurations).          |
| `configuration_id`         | string or null         | Saved configuration ID used for the job, if any.                                                                                 |
| `usage`                    | object or null         | Credits billed against the job. Requires `expand=usage`. See [Usage and credits](#usage-and-credits).                            |
| `metadata`                 | object or null         | Job-level metadata. None of the documented `expand` values return it, so expect `null`.                                          |

### Expanding the response

`GET /api/v2/extract/{job_id}` returns `extract_result` by default and omits the heavier fields. Add one or more `expand` values to include them. The create and cancel responses take no `expand` and always include `configuration`.

| `expand` value     | Adds                                                           |
| ------------------ | -------------------------------------------------------------- |
| `extract_metadata` | `extract_metadata` — per-field citations and confidence scores |
| `configuration`    | `configuration` — the resolved configuration the job ran with  |
| `usage`            | `usage` — credits billed against the job                       |

The list endpoint, `GET /api/v2/extract`, accepts `expand=configuration` and `expand=extract_metadata` and returns `items`, a `next_page_token` for the next page, and an optional `total_size`.

## Job status

| `status`    | Meaning                                         |
| ----------- | ----------------------------------------------- |
| `PENDING`   | Queued, not yet started.                        |
| `RUNNING`   | Actively processing.                            |
| `COMPLETED` | Finished; `extract_result` is populated.        |
| `FAILED`    | Terminated with an error; read `error_message`. |
| `CANCELLED` | Cancelled by the user.                          |

Poll until the status is one of the three terminal values, or let `client.extract.wait_for_completion(job.id)` do it for you.

## extract\_result

`extract_result` is the data itself, and its shape is set by two things: your `data_schema` decides the keys and value types, and `extraction_target` decides whether you get one object or an array of them.

| `extraction_target` | `extract_result` shape                                                                                   |
| ------------------- | -------------------------------------------------------------------------------------------------------- |
| `per_doc` (default) | A single JSON object matching your schema.                                                               |
| `per_page`          | An array of objects, one per page, each matching your schema.                                            |
| `per_table_row`     | An array of objects, one per detected entity (table row, list item, section), each matching your schema. |

A `per_doc` result for the schema in the [citations example](/llamaparse/extract/examples/extract_data_with_citations/index.md):

```
{
  "company_name": "NVIDIA Corporation",
  "filing_type": "10 K",
  "filing_date": "February 26, 2025",
  "fiscal_year": 2025,
  "unit": "millions",
  "revenue": 130497
}
```

### Repeating entities

With `per_table_row` the schema describes a single entity and the response is an array with one element per entity, so `len(job.extract_result)` is the entity count. The [repeating entities example](/llamaparse/extract/examples/extract_repeating_entities/index.md) returns this shape for a hospital directory:

```
[
  {"county": "Alameda", "hospital_name": "Alameda Hospital", "plan_names": ["Trio HMO", "SaveNet", "Access+ HMO"]},
  {"county": "Alameda", "hospital_name": "Eden Medical Center", "plan_names": ["Trio HMO", "Access+ HMO", "PPO"]}
]
```

Arrays declared inside your schema (for example `skills: list[str]` or `items: list[LineItem]`) appear as ordinary JSON arrays inside each result object, whatever the extraction target.

### Missing values

A field the document has no value for is handled according to your schema. An optional field (one not listed in `required`) is left out of the result. A required field, or one declared nullable, is returned as `null`. See [Required and optional fields](/llamaparse/extract/guides/configuring-extract/#required-and-optional-fields/index.md).

## extract\_metadata

`extract_metadata` is returned only with `expand=extract_metadata`. Its `field_metadata` entries carry citations and confidence scores only when the matching extension is turned on in the configuration. It has three fields:

| Field            | Type           | Description                                                        |
| ---------------- | -------------- | ------------------------------------------------------------------ |
| `field_metadata` | object or null | Per-field citations, confidence scores, and reasoning. See below.  |
| `parse_job_id`   | string or null | The Parse job that produced the document text the extraction read. |
| `parse_tier`     | string or null | Parse tier used for that Parse job.                                |

Turbo jobs produce no parse output; see [Tiers](/llamaparse/extract/guides/configuring-extract/#tiers/index.md). Jobs run with `spreadsheet_mode` produce neither citations nor confidence scores.

### field\_metadata

`field_metadata` mirrors `extract_result`, and which of its three keys is populated follows the extraction target:

| Key                 | Populated when                         | Shape                                                 |
| ------------------- | -------------------------------------- | ----------------------------------------------------- |
| `document_metadata` | `extraction_target` is `per_doc`       | One object keyed by field name.                       |
| `page_metadata`     | `extraction_target` is `per_page`      | An array of per-field metadata objects, one per page. |
| `row_metadata`      | `extraction_target` is `per_table_row` | An array of per-field metadata objects, one per row.  |

Inside each of those, the tree follows your schema: a scalar field maps to a metadata entry; an array field maps to a list where each element holds the entries for that element’s sub-fields, indexed by array position; a nested object holds entries for its sub-fields recursively. The metadata for `items[0].amount` therefore lives at `document_metadata.items[0].amount`, not at `document_metadata.items[0]`.

A metadata entry carries up to four keys, depending on which extensions are on:

| Key                     | Enabled by                | Description                                                           |
| ----------------------- | ------------------------- | --------------------------------------------------------------------- |
| `citation`              | `cite_sources: true`      | Array of source locations for the value. See [Citations](#citations). |
| `confidence`            | `confidence_scores: true` | Combined confidence, 0 to 1. The value to threshold on.               |
| `parsing_confidence`    | `confidence_scores: true` | How well the relevant context was parsed from the source document.    |
| `extraction_confidence` | `confidence_scores: true` | How well the extracted value matches the schema field.                |

The example from the API reference, with both extensions on, for a schema with `vendor`, `total`, and an `items` array:

```
"extract_metadata": {
  "field_metadata": {
    "document_metadata": {
      "vendor": {
        "citation": [
          {
            "page": 1,
            "matching_text": "Noisebridge",
            "bounding_boxes": [{"x": 72.0, "y": 96.0, "w": 120.0, "h": 14.0}]
          }
        ],
        "confidence": 1.0,
        "extraction_confidence": 1.0,
        "parsing_confidence": 1.0
      },
      "total": {
        "citation": [{"page": 1, "matching_text": "$10.00"}],
        "confidence": 1.0
      },
      "items": [
        {
          "amount": {"citation": [{"page": 1, "matching_text": "$10.00"}], "confidence": 1.0},
          "description": {"citation": [{"page": 1, "matching_text": "$10/month"}], "confidence": 0.998}
        }
      ]
    }
  }
}
```

### Citations

Set `cite_sources: true` in the configuration. Every leaf field then carries a `citation` array, with one element per place the value was found, so a field cited across several pages has several elements. Each element has:

| Key               | Type                    | Description                                                     |
| ----------------- | ----------------------- | --------------------------------------------------------------- |
| `page`            | integer                 | 1-based page number where the text was found.                   |
| `matching_text`   | string                  | The verbatim source text the value was extracted from.          |
| `bounding_boxes`  | array of `{x, y, w, h}` | Location of the cited text on the page.                         |
| `page_dimensions` | `{width, height}`       | Page size, for scaling the bounding boxes when you render them. |

On the `turbo` tier citations are text-level only: `page` and `matching_text` are returned without bounding boxes.

### Confidence scores

Set `confidence_scores: true` in the configuration. `confidence` combines `parsing_confidence` and `extraction_confidence` and is the score to route on; the other two explain where a low score came from. Calibration differs by tier, and long free-text fields score lower than short factual ones; see [Confidence scores](/llamaparse/extract/guides/extensions/#confidence-scores/index.md) for how to pick a threshold.

### Reasoning

Per-field reasoning strings appear in `field_metadata` for Extract versions through `2026-03-31` and are not returned by newer versions; see [Reasoning metadata](/llamaparse/extract/guides/extensions/#reasoning-metadata/index.md).

## Saved configurations

A job created with `configuration_id` runs with the saved configuration’s parameters, so the extensions, extraction target, and tier in that saved configuration decide what the response contains. The response records which configuration was used:

- `configuration_id` echoes the saved configuration ID. It is null for jobs created with an inline `configuration`.
- `configuration` (with `expand=configuration`) is the full resolved configuration the job ran with: `data_schema`, `tier`, `version`, `extraction_target`, `cite_sources`, `confidence_scores`, and the other options from [Configuring Extract](/llamaparse/extract/guides/configuring-extract/#configuration-options/index.md).
- `configuration.version` is the concrete release the job ran, reported as a release name such as `2.5` and fixed at job creation. It is never `latest`, even when the saved configuration stores `latest` or a date.

See [Using saved configurations](/llamaparse/extract/examples/using_saved_configurations/index.md) for creating and reusing them.

## Usage and credits

`usage`, returned with `expand=usage`, describes what the job consumed:

| Key               | Description                                                                                                                                                                            |
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `credits`         | Total credits billed against the job: the sum of the two components below, counting only those already recorded.                                                                       |
| `extract_credits` | Credits billed for the extraction itself.                                                                                                                                              |
| `parse_credits`   | Credits billed against the Parse job that Extract created for this job. `null` when `file_input` was a Parse job you created yourself, because those credits belong to that Parse job. |

`usage` is `null` until the job is `COMPLETED`, and each value inside it reads `null` until billing has recorded it, which can trail job completion, so poll again rather than reading `null` as zero. One Parse job can back several Extract jobs, and each of them reports that same `parse_credits`; to total across jobs, sum `extract_credits` and count the parse once. See [Check the credits a job billed](/llamaparse/extract/api/#5-check-the-credits-a-job-billed/index.md).

## Fetch the full response

```
import os
from llama_cloud import LlamaCloud


client = LlamaCloud(api_key=os.environ["LLAMA_CLOUD_API_KEY"])


DATA_SCHEMA = {
    "type": "object",
    "properties": {
        "company_name": {"type": "string", "description": "Name of the company"},
        "revenue": {"type": "number", "description": "Annual revenue in USD"},
    },
}


file_obj = client.files.create(file="path/to/document.pdf", purpose="extract")


job = client.extract.create(
    file_input=file_obj.id,
    configuration={
        "data_schema": DATA_SCHEMA,
        "tier": "agentic",
        "cite_sources": True,
        "confidence_scores": True,
    },
)
job = client.extract.wait_for_completion(job.id)


# Metadata, configuration, and usage are omitted unless expanded
job = client.extract.get(job.id, expand=["extract_metadata", "configuration", "usage"])


print(job.status)
print(job.extract_result)
print(job.extract_metadata.field_metadata.document_metadata)
print(job.usage)
```
