Skip to content
Guide
Parse
Features

Charts and figures

How Parse turns charts into structured table data with specialized chart parsing, and how to get figures, embedded images, and page screenshots out of a document as downloadable image files.

Parse handles charts and figures two ways. Specialized chart parsing reads a bar, line, or pie chart and returns its values as a table in the items tree, so a model can reason over the numbers instead of the pixels. Image output saves the figures themselves (cropped layout regions, embedded images, or full-page screenshots) as files you download by presigned URL.

  • Earnings decks, scientific papers, and dashboards where the data you need lives only in a chart.
  • Multimodal pipelines that want one markdown blob plus a screenshot of every page.
  • Figures and diagrams you want to show in a UI or pass to a vision model separately from the text.
  • Self-contained markdown that carries its images inline.

Specialized chart parsing is on by default at the agentic_plus tier and opt-in on cost_effective and agentic. The fast tier runs no AI model, so it does not extract chart data.

OptionTypeDefaultWhat it does
processing_options.specialized_chart_parsing"efficient", "agentic", or "agentic_plus"unset (on by default for agentic_plus)Extract chart data as structured tables in the items tree. Any of the three values turns it on. Retrieve with expand=["items"].
output_options.images_to_savearray of "screenshot", "embedded", "layout"saves layout when the output links to cropped imagesWhich image files to save: full-page renders, images found in the document, or cropped figures and diagrams. Pass [] to save none. Retrieve with expand=["images_content_metadata"].
output_options.markdown.inline_imagesbooleanunsetEmbed images in the markdown itself instead of referencing separately saved files.
processing_options.ignore.ignore_text_in_imagebooleanunsetSkip OCR text from embedded images when it is noise, such as logos. Turns OCR off for the job, so born-digital files only.
input_options.presentation.skip_embedded_databooleanunsetFor PPTX and Keynote, skip extracting the data behind native charts and keep only their visual.

The image_filenames query parameter on the result endpoint narrows images_content_metadata to the files you name.

Parse with chart parsing on, save the cropped figures, and pull the first table off the chart page:

from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse(
file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse")
tier="agentic_plus",
version="latest",
processing_options={"specialized_chart_parsing": "agentic_plus"},
output_options={"images_to_save": ["layout"]},
expand=["items", "images_content_metadata"],
)
page = next(p for p in result.items.pages if p.success and p.page_number == 3) # the chart page
tables = [item for item in page.items if item.type == "table"]
if tables:
for row in tables[0].rows:
print(row)
for image in result.images_content_metadata.images:
print(f"{image.filename}: {image.presigned_url}")

With chart parsing, Parse often represents a chart’s data as a table item on that page. For the grouped bar chart in the chart example, the first row is the header and the rest are the series values:

['Fiscal Year', 'Budget Deficit (Billions of Dollars)', 'Net Operating Cost (Billions of Dollars)']
['2020', '$3,131.9', '$3,841.4']
['2021', '$2,775.6', '$3,094.9']
['2022', '$1,375.5', '$4,171.0']

Figures that stay as images appear in the items tree as image items, and images_content_metadata lists every saved file with a download URL:

{
"images_content_metadata": {
"total_count": 3,
"images": [
{ "index": 0, "filename": "image_0.png", "category": "layout", "content_type": "image/png", "presigned_url": "https://..." }
]
}
}

Presigned URLs expire. Download promptly, or call client.parsing.get(job_id=..., expand=["images_content_metadata"]) again for fresh ones; see fetch specific images to pull a subset by file name. For full-page screenshots, request expand=["metadata"] too and use each page’s original_orientation_angle when overlaying boxes; see Layout and bounding boxes.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/