Parse response format
Field-by-field reference for what a Parse job returns — the job object, per-page markdown, text, items, metadata and forms, item types, bounding-box conventions, the grounded-items sidecar, image metadata, and download URLs.
A Parse job result is one JSON object. job is always present; every other field appears only when you ask for it with expand. This page is the reference for what each field contains. For which expand values to combine, how to poll, and the common retrieval patterns with code, see Retrieving Results. For the request options that produce each output, see Configuring Parse.
Response envelope
Section titled “Response envelope”{ "job": { "id": "JOB_ID", "status": "COMPLETED", "...": "..." }, "markdown": { "pages": [ ... ] }, "markdown_full": "# Complete Document\n\n...", "text": { "pages": [ ... ] }, "text_full": "Complete Document\n\n...", "items": { "pages": [ ... ] }, "metadata": { "pages": [ ... ], "document": { ... } }, "forms": { "pages": [ ... ] }, "job_metadata": { ... }, "images_content_metadata": { "total_count": 3, "images": [ ... ] }, "result_content_metadata": { "markdown": { ... }, "grounded_items": { ... } }}| Field | Type | Filled by expand | Contents |
|---|---|---|---|
job | object | always | Job status and identifiers. See Job object. |
markdown | object | markdown | Per-page markdown. See Markdown pages. |
markdown_full | string | markdown_full | The whole document as one markdown string. |
text | object | text | Per-page plain text. See Text pages. |
text_full | string | text_full | The whole document as one plain-text string. |
items | object | items | Per-page structured items. See Items pages. |
metadata | object | metadata | Per-page and document-level metadata. See Metadata. |
forms | object | forms | Per-page form trees. See Forms (Beta). |
job_metadata | object | job_metadata | Job execution metadata: state-transition timestamps and timing. Not the job configuration. |
images_content_metadata | object | images_content_metadata | Saved images with download URLs. See Images. |
result_content_metadata | object | any *_content_metadata value | Download URLs for result files. See Download URLs. |
Fields you did not request are null or absent. Two tier rules apply: forms is never populated on the fast tier, because the form pass does not run there, and a fast job pinned to a version older than 2026-06-15 produces text only, so expand=markdown is rejected with 400 and items comes back null. See Tiers.
Job object
Section titled “Job object”job is present on every response, including while the job is still running.
| Field | Type | Meaning |
|---|---|---|
id | string | Parse job identifier. Pass it to later get calls. |
status | string | PENDING, RUNNING, COMPLETED, FAILED or CANCELLED. Content fields are populated only once the status is COMPLETED. |
project_id | string | Project the job belongs to. |
name | string or null | Optional display name. |
tier | string or null | Tier the job ran with (fast, cost_effective, agentic, agentic_plus). |
created_at, updated_at | string or null | Timestamps. |
error_message | string or null | Error details when status is FAILED. |
usage | object or null | { "credits": 30.0 }. Requires expand=usage. credits is null until billing has recorded the job, which can trail completion by a short time. |
user_metadata | object or null | The key/value tags you attached to the request, returned verbatim. |
Page arrays
Section titled “Page arrays”Every per-page representation has the same outer shape: a pages array with one entry per parsed page, in document order. page_number is 1-based and refers to the page’s position in the source document, so it is the key to join markdown, items, metadata and forms entries for the same page, and to match a page screenshot.
| Representation | Page carries success | Page carries page_width / page_height |
|---|---|---|
markdown.pages[] | yes | no |
text.pages[] | no | no |
items.pages[] | yes | yes |
metadata.pages[] | no | no |
forms.pages[] | yes | yes |
Where a page carries success, a page that could not be processed is replaced by a failed-page entry with the same page_number and none of the content fields:
{ "page_number": 2, "success": false, "error": "..." }Check success before reading content fields on those pages. In the typed SDKs the page is a union of the success and failure shapes, so the check also narrows the type. Failed pages count toward the job’s allowed_page_failure_ratio (see timeouts and failure conditions).
Markdown pages
Section titled “Markdown pages”expand=markdown returns one entry per page. markdown_full returns the same content joined into a single string instead.
| Field | Type | Meaning |
|---|---|---|
page_number | integer | 1-based page number. |
markdown | string | Page content as markdown. Tables render as HTML <table> blocks or markdown pipe tables depending on output_options.markdown.tables.output_tables_as_markdown. |
header | string or null | Page header text, in markdown, when one was detected. |
footer | string or null | Page footer text, in markdown, when one was detected. |
line_numbers | array or null | Printed gutter line numbers mapped to offsets in markdown. Only with output_options.markdown.annotate_line_numbers. |
success | true | Always true on a successful page. |
{ "page_number": 1, "success": true, "markdown": "# Heading\n\n## Subheading\n\nContent with **formatting**...", "header": "Page header", "footer": "LlamaIndex 2026", "line_numbers": [ { "line_number": "22", "start_index": 0, "end_index": 34 } ]}Printed line numbers
Section titled “Printed line numbers”When output_options.markdown.annotate_line_numbers is on, Parse detects physical left-gutter line numbers, removes the printed labels from the content, and maps each one to the text it labelled:
| Field | Type | Meaning |
|---|---|---|
line_number | string | The printed value, kept as a string because documents print 22 and A-3 alike. |
start_index | integer | Zero-based, inclusive offset into that page’s final markdown. |
end_index | integer | Zero-based, exclusive offset. |
Offsets count UTF-16 code units, matching JavaScript String.slice. The array is omitted when no printed line can be mapped confidently.
Text pages
Section titled “Text pages”expand=text returns the plain-text layer of each page with no markup. With spatial text options set, whitespace preserves the visual layout.
| Field | Type | Meaning |
|---|---|---|
page_number | integer | 1-based page number. |
text | string | Plain text of the page. |
{ "text": { "pages": [ { "page_number": 1, "text": "Extracted plain text content..." } ] } }Items pages
Section titled “Items pages”expand=items returns the layout tree: every element Parse detected on the page, typed and in reading order. Use it when you need tables as data, figures as files, or the position of an element on the page.
| Field | Type | Meaning |
|---|---|---|
page_number | integer | 1-based page number. |
page_width, page_height | number | Page size in points. Bounding boxes on this page are in the same units. |
items | array | The item objects below, in reading order. |
revisions | array or null | Word tracked changes and comments. Only with output_options.markdown.annotate_revisions. See Revisions. |
success | true | Always true on a successful page. |
{ "page_number": 1, "page_width": 612.0, "page_height": 792.0, "success": true, "items": [ { "type": "heading", "level": 1, "value": "Document Title", "md": "# Document Title", "bbox": [ { "x": 72.0, "y": 60.0, "w": 300.0, "h": 24.0 } ] }, { "type": "table", "rows": [["Header1", "Header2"], ["Row1", "Data1"]], "html": "<table>...</table>", "csv": "Header1,Header2\nRow1,Data1", "md": "| Header1 | Header2 |\n|---|---|..." } ]}Item types
Section titled “Item types”Every item carries type, md (its markdown rendering) and bbox (an array of bounding boxes, or null). The other fields depend on the type.
type | Represents | Type-specific fields |
|---|---|---|
text | A paragraph or other run of body text | value (plain text) |
heading | A heading | level (1–6), value |
table | A table, including chart data when chart parsing is on | rows, html, csv, merged_from_pages, merged_into_page, parse_concerns — see Tables |
image | A figure, photo or diagram | url (link to the extracted image), caption |
list | An ordered or unordered list | ordered (boolean), items (nested text and list items) |
code | A code block | value (the code), language (identifier or null) |
link | A hyperlink | text (display text), url |
header | The page’s running header | items (the items inside it) |
footer | The page’s running footer | items (the items inside it) |
header and footer are containers: their items array holds the headings, text, images and so on that sit in the header or footer region, each a normal item. The markdown page object exposes the same regions as plain strings in its header and footer fields.
Tables
Section titled “Tables”| Field | Type | Meaning |
|---|---|---|
rows | array of arrays | Cell values row by row. Each cell is a string, a number or null. |
html | string | The table as an HTML <table>. Keeps merged cells (colspan / rowspan). |
csv | string | The table as CSV. |
md | string | The table as a markdown pipe table. Cannot express merged cells. |
merged_from_pages | array of integers or null | When merge_continued_tables is on, the page numbers whose table fragments were merged into this one. |
merged_into_page | integer or null | On a fragment that was merged elsewhere: the page where the full merged table starts. |
parse_concerns | array or null | Quality flags from table extraction. Each entry has type (for example inconsistent_row_cell_count) and human-readable details. |
With specialized chart parsing enabled, the data read from a chart is returned as a table item on that page, so the series values are available in rows like any other table. The chart parsing example walks through loading one into a DataFrame.
Revisions
Section titled “Revisions”When output_options.markdown.annotate_revisions is on, an items page that contains Word tracked changes or reviewer comments also carries a revisions array:
{ "page_number": 1, "items": [ ... ], "revisions": [ { "type": "deleted", "target": "within thirty (30) days", "content": "within thirty (30) days", "author": "J. Reviewer", "target_bbox": { "x": 120.4, "y": 302.1, "w": 96.2, "h": 11.0 }, "revision_bbox": { "x": 460.0, "y": 298.5, "w": 140.0, "h": 24.0 }, "start_index": 512, "end_index": 536 } ], "success": true}| Field | Type | Meaning |
|---|---|---|
type | string | One of inserted, deleted, formatted, moved_from, moved_to, comment. |
target | string | The page text the revision applies to. |
content | string | The revision or comment content. |
author | string or null | Reviewer name, when available. |
target_bbox | object | Bounding box of the target text: x, y, w, h in page points. |
revision_bbox | object | Bounding box of the printed revision balloon. |
target_spans | array or null | Present when the target is discontinuous: each span has its own target, target_bbox, start_index and end_index. |
start_index, end_index | integer or null | Offsets of the target in that page’s final markdown, inclusive start and exclusive end. |
Bounding boxes
Section titled “Bounding boxes”Items, form fields, revisions and images carry bounding boxes that locate them on the page. On items and form fields bbox is an array, because one element can occupy several rectangles: a paragraph that wraps around a figure, or a form field whose fillable area is split.
| Field | Type | Meaning |
|---|---|---|
x, y | number | Position of the box’s top-left corner. |
w, h | number | Width and height. |
r | number, optional | Clockwise rotation of the visible text in degrees, around the box’s center. Omitted when unrotated. |
label | string or null | Label for the box, when one applies. |
confidence | number or null | Confidence score for the box, when available. |
start_index, end_index | integer or null | Optional character offsets into the item’s text. |
Conventions, the same for every box type:
- Units are page points, not pixels and not normalized. The page’s
page_widthandpage_height(on the items page, forms page and grounded-items sidecar) give the frame the box lives in. A US Letter page is612 × 792. - Origin is the top-left corner of the page,
xincreasing to the right andyincreasing downward, matching image coordinates. - Boxes are relative to the page the item sits on. Read
page_widthandpage_heightfrom that page entry; pages in one document can differ in size. x/y/w/hdescribe the unrotated rectangle;rrotates it. Parse has already composed the page’soriginal_orientation_angle(frommetadata) into the box. Apply onlyr; never add the page orientation to it.
Worked example. On a 612 × 792 page, a word box of { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } sits 72 points from the left edge and 100 points from the top. To draw it on a page screenshot of 1224 × 1584 pixels, scale by 1224 / 612 = 2 horizontally and 1584 / 792 = 2 vertically: the rectangle runs from pixel (144, 200) to (214, 224). When the two scale factors differ, rotate the corners by r first and scale after, or the box distorts. The granular bounding boxes example has a complete rendering helper that handles rotation and page orientation.
Camera photos. For photos corrected with camera_photo_correction, Parse maps each rotated text rectangle back into the original photo frame and exports the best oriented-rectangle approximation. A perspective quadrilateral cannot be represented exactly by one rectangle plus one angle, so small edge differences are possible on strongly skewed photos. If that approximation would extend beyond the photo boundary, Parse clips it to a bounded axis-aligned box and sets r to 0.
Grounded items sidecar
Section titled “Grounded items sidecar”Item-level boxes are the finest grain inlined on the result. Setting output_options.granular_bboxes to any of "word", "line", "cell" makes Parse write a separate grounded-items file with per-line, per-word or per-table-cell boxes for every item. Its download URL is included automatically as result_content_metadata.grounded_items; nothing extra goes in expand.
The file is JSONL, one JSON object per line and one line per page, not a JSON array. Each row is either a success row or a failure row:
// Success row{ "page_number": 1, "page_width": 612, "page_height": 792, "success": true, "items": [ { "type": "text", "md": "Hello world", "bbox": [{ "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }], "grounding": { "source": "md", "lines": [ { "span": [0, 11], "bbox": { "x": 72.0, "y": 100.0, "w": 200.0, "h": 12.0 }, "words": [ { "span": [0, 5], "bbox": { "x": 72.0, "y": 100.0, "w": 35.0, "h": 12.0 } }, { "span": [6, 11], "bbox": { "x": 110.0, "y": 100.0, "w": 40.0, "h": 12.0 } } ] } ] } } ]}
// Failure row — grounding could not be produced for this page{ "page_number": 2, "success": false, "error": "..." }| Field | Meaning |
|---|---|
items[].type, md, bbox | The same item fields as the inline items representation. Nested items (list entries) appear under items[].items. |
items[].grounding.source | Which text the spans index into: md for the item’s markdown, caption for an image caption. |
items[].grounding.lines[] | One entry per text line: a bbox and a [start, end) span into the source text. |
items[].grounding.lines[].words[] | One entry per word, when "word" was requested: a bbox and a span. |
items[].grounding.rows[row][col] | For table items: per-cell bbox, span and lines, with null for missing cells, plus row_bboxes and column_bboxes. |
Spans are UTF-8 byte offsets, not character offsets. Encode the source text as UTF-8 before slicing, or the offsets drift on any non-ASCII character:
md_bytes = item["md"].encode("utf-8")for line in item["grounding"]["lines"]: start, end = line["span"] print(md_bytes[start:end].decode("utf-8"), line["bbox"])For a complete worked example, from request to walking the per-word grounding, see the granular bounding boxes example.
Metadata
Section titled “Metadata”expand=metadata returns per-page quality and provenance fields plus a document block.
Page metadata
Section titled “Page metadata”| Field | Type | Meaning |
|---|---|---|
page_number | integer | 1-based page number. |
confidence | number or null | How confident Parse is in this page’s output, 0–1. With confidence_score_effort: "high", this reflects the high-effort assessment. |
cost_optimized | boolean or null | true when Cost Optimizer routed this page to cost_effective. |
triggered_auto_mode | boolean or null | true when an auto_mode_configuration rule fired on this page. |
original_orientation_angle | integer or null | Rotation Parse applied to read the page, in degrees; 0 when none was needed. Already composed into every bounding box on the page. |
printed_page_number | string or null | The page number as printed on the page ("i", "A-3"). Only with output_options.extract_printed_page_number. |
watermark | string or null | Watermark text detected on the page; several are joined with |. Absent when none. Version 2026-09-28 or later on cost_effective, agentic and agentic_plus, whatever output_options.watermark_handling is set to; see Watermarks. |
speaker_notes | string or null | Presentation inputs only: the slide’s speaker notes. |
slide_section_name | string or null | Presentation inputs only: the section the slide belongs to. |
{ "page_number": 1, "confidence": 0.985, "cost_optimized": false, "triggered_auto_mode": false, "original_orientation_angle": 0, "printed_page_number": null, "watermark": "CONFIDENTIAL", "speaker_notes": null, "slide_section_name": null}Document metadata
Section titled “Document metadata”metadata.document is populated when confidence_score_effort: "high" is set.
| Field | Type | Meaning |
|---|---|---|
confidence | number or null | Mean confidence across the pages the high-effort judge scored, 0–1. |
confidence_breakdown.min_page_score | number | The lowest page score, the worst page in the document. |
confidence_breakdown.scored_pages | integer | Pages the judge scored. |
confidence_breakdown.total_pages | integer | Pages in the document. |
Forms (Beta)
Section titled “Forms (Beta)”expand=forms is populated only on jobs created with processing_options.forms: "enrich"; otherwise it is null. Every page of the document has an entry; a page with no detected form has "forms": [].
| Field | Type | Meaning |
|---|---|---|
page_number | integer | 1-based page number. |
page_width, page_height | number or null | Page size in points, the frame for field bbox values. |
detected_form_types | array of strings or null | Form types detected on the page; null when the page was not treated as a form. |
forms | array | One object per form detected on the page. |
success | true | Always true on a successful page. |
Each form holds the same content twice: json, an ordered tree of nodes, and list, a flattened bullet list whose md drops straight into a prompt.
Node type | Represents | Fields |
|---|---|---|
section | A grouping printed on the form (Part III, box 15) | id, label, items (child nodes in reading order) |
field | One entry: text input, checkbox, select group or signature line | field, id, label, value, isEmpty, valueItems, bbox |
table | A fillable grid | id, label, columns, rows, bbox |
How a field node reads depends on its field kind:
field | value | Notes |
|---|---|---|
text | The entered text, verbatim | A printed-but-blank field has isEmpty: true and no value. |
checkbox | boolean, checked or not | |
signature | boolean, signed or not | |
single_select, multi_select | none | valueItems lists the options, usually checkbox fields with their own boolean value. |
bbox on a field or table is an array of bounding boxes around the fillable area, in page points. In a table node, each cell of rows is a string, null for a printed-but-blank cell, or { "items": [...] } holding the cell’s own nodes (a checkbox column, for example). When the job also sets output_options.granular_bboxes, a node may carry an optional grounding object that locates its printed id, label and string value (and, for tables, columns and scalar cells) as per-line boxes with UTF-8 byte spans into that text.
{ "page_number": 1, "page_width": 612, "page_height": 792, "success": true, "forms": [ { "json": [ { "type": "field", "field": "text", "id": "1", "label": "Wages, tips, other compensation", "value": "29,513", "bbox": [ { "x": 349.2, "y": 96.5, "w": 114.0, "h": 12.2 } ] }, { "type": "field", "field": "multi_select", "id": "13", "valueItems": [ { "type": "field", "field": "checkbox", "label": "Statutory employee", "value": true }, { "type": "field", "field": "checkbox", "label": "Retirement plan", "value": false } ] } ], "list": { "type": "list", "ordered": false, "md": "- [1] Wages, tips, other compensation: 29,513\n- [13]\n - [x] Statutory employee\n - [ ] Retirement plan", "items": [ ... ] } } ]}The Python SDK exposes json as json_, valueItems as value_items and isEmpty as is_empty; the raw response and the TypeScript SDK use the names shown here. For a walker that reaches every field through nested sections and table cells, see the enriched forms example.
Images
Section titled “Images”expand=images_content_metadata lists the images saved for the job, each with its own download URL. Which images exist depends on output_options.images_to_save.
{ "images_content_metadata": { "total_count": 2, "images": [ { "index": 0, "filename": "image_0.png", "category": "screenshot", "content_type": "image/png", "bbox": { "x": 0, "y": 0, "w": 612, "h": 792 }, "presigned_url": "https://..." }, { "index": 1, "filename": "image_1.jpg", "category": "layout", "content_type": "image/jpeg", "bbox": { "x": 72, "y": 300, "w": 468, "h": 210 }, "presigned_url": "https://..." } ] }}| Field | Type | Meaning |
|---|---|---|
total_count | integer | Number of images returned. |
images[].index | integer | Position in extraction order. |
images[].filename | string | File name, for example image_0.png. Pass it to the image_filenames parameter to fetch a subset. |
images[].category | string or null | screenshot (a full-page render), embedded (an image file inside the document) or layout (a figure cropped from the page). |
images[].content_type | string or null | MIME type. |
images[].bbox | object or null | Where the image sits on its page: x, y, w, h in page coordinates, rounded to integers. |
images[].presigned_url | string or null | Temporary download URL. |
images[].size_bytes | — | Deprecated; always null. |
image items in the items tree carry their own url to the image they represent.
Rotated pages. A full-page screenshot is rendered in the page’s original orientation. To draw bounding boxes on it, also request expand=metadata and read original_orientation_angle for the same page_number (match on page number, not on image index). Either rotate the screenshot upright first, or apply the inverse page rotation to the box corners; apply only each box’s own r on top. The rendering recipe shows both.
Download URLs
Section titled “Download URLs”Every *_content_metadata value adds an entry under result_content_metadata instead of returning the content inline. Each entry has the same shape:
{ "result_content_metadata": { "markdown": { "size_bytes": 45678, "exists": true, "presigned_url": "https://..." }, "items": { "size_bytes": 67890, "exists": true, "presigned_url": "https://..." }, "grounded_items": { "size_bytes": 123456, "exists": true, "presigned_url": "https://..." } }}| Field | Meaning |
|---|---|
size_bytes | Size of the result file. Check it before deciding whether to download. |
exists | Whether the file was produced for this job. |
presigned_url | Temporary download URL. Request the result again with the same expand to get a fresh one once it expires. |
The key under result_content_metadata and the file you download:
expand value | Key | File |
|---|---|---|
markdown_content_metadata | markdown | Per-page markdown, JSON with the same pages shape as the inline field |
text_content_metadata | text | Per-page text, JSON |
items_content_metadata | items | Items tree, JSON |
metadata_content_metadata | metadata | Page metadata, JSON |
forms_content_metadata | forms | Forms, JSON. Jobs with processing_options.forms: "enrich" only |
markdown_full_content_metadata | markdown_full_content_metadata | The whole document as one .md file |
text_full_content_metadata | text_full_content_metadata | The whole document as one .txt file |
xlsx_content_metadata | xlsx | Tables as a workbook. Jobs with tables_as_spreadsheet.enable: true only |
output_pdf_content_metadata | outputPDF | The rendered PDF. Jobs with output_options.save_output_pdf: true only |
raw_words_content_metadata | raw_words | One JSON object per word with page number and box, as JSONL. Jobs with output_options.additional_outputs: ["word_bbox"] only |
| none needed | grounded_items | The grounded items sidecar, JSONL. Jobs with output_options.granular_bboxes set only |
images_content_metadata is the exception: it returns the images object at the top level rather than an entry here.
Results are not paginated. A per-page representation comes back as the complete pages array in one response, and markdown_full and text_full as one string. For long documents, request the *_content_metadata variant, read size_bytes, and download the file instead of holding the content in the response.
See also
Section titled “See also”- Retrieving Results — which
expandvalues to combine, polling, and the common patterns with code in every SDK - Configuring Parse — the request options that produce each output
- Parse a PDF and interpret outputs — the four main views on a real document
- Granular bounding boxes — word, line and cell grounding, with a rendering helper
- Enriched forms output — walking the form tree
- Parse API reference — the generated schema listing