Skip to content
Guide
For Agents
Integrations

Use LlamaParse with LangChain

Turn Parse output into LangChain Document objects for splitting, embedding and retrieval, using the current llama-cloud SDK.

LlamaParse is framework-agnostic: it returns Markdown, text or JSON over the API, and any framework that works with text can consume it. For LangChain, the shortest path is to run a Parse job with the llama-cloud SDK and wrap each page in a LangChain Document.

Terminal window
pip install langchain-core langchain-text-splitters "llama-cloud>=2.8"
export LLAMA_CLOUD_API_KEY=llx-...
from langchain_core.documents import Document
from llama_cloud import LlamaCloud
client = LlamaCloud() # reads LLAMA_CLOUD_API_KEY
file = client.files.create(file="data/report.pdf", purpose="parse")
result = client.parsing.parse(
file_id=file.id, tier="agentic", version="latest", expand=["markdown"]
)
documents = [
Document(
page_content=p.markdown,
metadata={"source": "data/report.pdf", "page": p.page_number},
)
for p in result.markdown.pages
if p.success
]

Pages that failed to parse come back with success set to False and no Markdown, so the comprehension skips them. From here the documents go through the usual LangChain steps, for example a text splitter and a vector store:

from langchain_text_splitters import MarkdownHeaderTextSplitter
splitter = MarkdownHeaderTextSplitter(
headers_to_split_on=[("#", "h1"), ("##", "h2")]
)
chunks = [
chunk
for doc in documents
for chunk in splitter.split_text(doc.page_content)
]

Markdown is the output to prefer here: Parse preserves headings and tables as Markdown structure, which header-aware splitters use to keep sections together.

Extract returns JSON in your schema, which maps directly onto a LangChain tool or a structured-output step. Run the job with client.extract.run(file_input=FILE_ID, configuration={"data_schema": DATA_SCHEMA}), which waits for completion, and pass job.extract_result to the chain. See Getting started with Extract for the schema options.

Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/