Tables
How Parse extracts tables from PDFs, scans, and images into markdown, HTML, structured rows, CSV, or an XLSX file, with the options for merged cells, multi-page tables, and borderless tables.
Parse recovers tables with their cell structure intact and returns each one in several forms at once: as part of the page markdown (pipe table or HTML), and as a typed table item in the items tree with rows, csv, html, and md representations. You choose the format with a few output options, and you can also have Parse write every table into an XLSX workbook.
When to use it
Section titled “When to use it”- Financial reports, 10-Ks, earnings decks, and audit reports with many tables across many pages.
- Tables with merged header cells, which markdown pipe tables cannot represent.
- Tables that continue across page breaks and should come back as one table.
- Borderless or lightly formatted tables that a text-based parser misses.
- Loading table data into pandas, a warehouse, or a spreadsheet rather than reading it as prose.
Pick cost_effective for mostly-text documents with simple tables, agentic for real tables and scanned pages, and agentic_plus for dense financial reports and complex tables. The fast tier runs no AI model, so pick cost_effective or higher when tables have merged cells, continue across pages, or lack borders.
Options
Section titled “Options”| Option | Type | Default | What it does |
|---|---|---|---|
output_options.markdown.tables.output_tables_as_markdown | boolean | unset | true emits markdown pipe tables; false keeps HTML <table> tags. Pipe tables are simpler but cannot represent merged cells; HTML can, with colspan. |
output_options.markdown.tables.merge_continued_tables | boolean | unset | Merge a table that spans several pages into one. The merged table appears on the first page with merged_from_pages metadata. |
output_options.markdown.tables.compact_markdown_tables | boolean | unset | Remove whitespace padding inside markdown table cells. |
output_options.markdown.tables.markdown_table_multiline_separator | string | unset | Separator for multi-line cell content in markdown tables, for example "<br>" to keep line breaks or " " to join. |
output_options.tables_as_spreadsheet.enable | boolean | unset | Also write every table into an XLSX file, one sheet per table. Retrieve with expand=["xlsx_content_metadata"]. |
output_options.tables_as_spreadsheet.guess_sheet_name | boolean | true | Name each sheet from the table’s headers and surrounding text instead of Table_1. |
processing_options.aggressive_table_extraction | boolean | unset | Try harder to find table boundaries, including tables without visible borders. May add false positives. |
processing_options.disable_heuristics | boolean | unset | Turn off outlined-table extraction and adaptive long-table handling when they produce wrong results. |
output_options.granular_bboxes | array of "cell", "line", "word" | [] | Add a bounding box per table cell for highlighting. See Layout and bounding boxes. |
Example
Section titled “Example”Parse a report, collect every table from the items tree, and load each one into pandas from its csv:
import io
import pandas as pdfrom llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY from the environment
result = client.parsing.parse( file_id="FILE_ID", # uploaded with client.files.create(file=..., purpose="parse") tier="agentic", version="latest", output_options={"markdown": {"tables": {"merge_continued_tables": True}}}, expand=["markdown", "items"],)
tables = [ (page.page_number, item) for page in result.items.pages if page.success # a failed page carries no items for item in page.items if item.type == "table"]print(f"Found {len(tables)} tables across {len(result.items.pages)} pages")
for page_number, table in tables: df = pd.read_csv(io.StringIO(table.csv)) print(f"page {page_number}: {len(df)} rows x {len(df.columns)} cols")What you get
Section titled “What you get”In the page markdown, a default agentic parse of the Quick Start report returns the summary table as HTML, keeping its merged header row:
<table> <thead> <tr><th colspan="3">Financial Measures (Dollars in Billions):</th></tr> <tr><th></th><th>2024</th><th>2023*</th></tr> </thead> <tbody> <tr><td>Gross Costs</td><td>$ (7,772.2)</td><td>$ (7,661.7)</td></tr> <tr><td>Less: Earned Revenue</td><td>$ 652.9</td><td>$ 539.5</td></tr> </tbody></table>In the items tree, the same table is a typed item with the data in four forms:
{ "type": "table", "rows": [["Financial Measures (Dollars in Billions):", "2024", "2023*"], ["Gross Costs", "$ (7,772.2)", "$ (7,661.7)"]], "csv": "Financial Measures (Dollars in Billions):,2024,2023*\nGross Costs,$ (7,772.2),$ (7,661.7)", "html": "<table>...</table>", "md": "| Financial Measures (Dollars in Billions): | 2024 | 2023* |\n|---|---|---|..."}rows is a list of lists, csv loads straight into pandas.read_csv, and html and md are ready to display. With tables_as_spreadsheet enabled, expand=["xlsx_content_metadata"] adds a presigned download URL for the workbook.
See also
Section titled “See also”- Parse a financial report and extract every table: the full walk-through in Python, TypeScript, Go, Java, and the CLI
- Parse charts in PDFs and analyze with pandas: the single-table variant, plus chart data as tables
- Configuring Parse: markdown output options and tables as spreadsheet
- Response format: tables for every field on a table item
- Spreadsheets for XLSX and CSV inputs, the other direction