---
title: Use LlamaParse from the LlamaIndex Framework | Developer Documentation
description: Feed Parse output into a LlamaIndex Framework index or agent, and use Extract results as structured tool output, with the current llama-cloud SDK.
---

The [LlamaIndex Framework](/python/framework/index.md) is our open-source Python toolkit for building agents and retrieval-augmented generation (RAG) over your own data. Its built-in readers handle clean text well. For scanned PDFs, forms, spreadsheets and slide decks, run a Parse job first and hand the result to the framework: the quality of what goes into an index decides the quality of every answer it gives.

Both products use the same API key. Set `LLAMA_CLOUD_API_KEY` in your environment and install the two packages:

Terminal window

```
pip install llama-index "llama-cloud>=2.8"
export LLAMA_CLOUD_API_KEY=llx-...
```

## Build an index from Parse output

`parsing.parse()` uploads the file, waits for the job to finish, and returns the result. With `expand=["markdown"]` the result carries one Markdown string per page. Wrap each page in a framework `Document` and build an index as you would from any other text:

```
from llama_cloud import LlamaCloud
from llama_index.core import Document, VectorStoreIndex


client = LlamaCloud()  # reads LLAMA_CLOUD_API_KEY


file = client.files.create(file="data/report.pdf", purpose="parse")
result = client.parsing.parse(
    file_id=file.id, tier="agentic", version="latest", expand=["markdown"]
)
pages = result.markdown.pages
documents = [
    Document(text=p.markdown, metadata={"page": p.page_number})
    for p in pages
    if p.success
]


index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("What does this document say about pricing?")
print(response)
```

Pages that failed to parse come back with `success` set to `False` and no Markdown, which is why the comprehension filters on it. Keeping the page number in `metadata` lets retrieval results point back to the page they came from.

The same pattern works for a folder of files: run one Parse job per file, collect the documents, and build a single index. For larger collections, [Index](/llamaparse/cloud-index-v2/getting_started/index.md) keeps an index in sync with your data sources and serves retrieval over the API, so nothing needs to be rebuilt locally.

## Use Extract output in an agent

Extract returns JSON shaped by your schema. A framework agent can call it as a tool, so the agent decides when to pull structured data out of a document:

```
from llama_cloud import LlamaCloud
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI


client = LlamaCloud()




def extract_invoice(path: str) -> dict:
    """Extract vendor, invoice number, date and total from an invoice file."""
    file = client.files.create(file=path, purpose="extract")
    job = client.extract.run(  # creates the job and waits for it to finish
        file_input=file.id,
        configuration={
            "data_schema": {
                "type": "object",
                "properties": {
                    "vendor": {"type": "string"},
                    "invoice_number": {"type": "string"},
                    "date": {"type": "string"},
                    "total": {"type": "number"},
                },
            }
        },
    )
    return job.extract_result




agent = FunctionAgent(
    tools=[extract_invoice],
    llm=OpenAI(model="gpt-4.1-mini"),
    system_prompt="You answer questions about invoices by extracting their fields.",
)
```

See [Getting started with Extract](/llamaparse/extract/sdk/index.md) for schema options and [Building agents](/python/framework/understanding/agent/index.md) for the agent side.

## Choosing where parsing runs

| Need                                                  | Use                                                          |
| ----------------------------------------------------- | ------------------------------------------------------------ |
| Clean text files, Markdown, HTML                      | The framework’s built-in readers                             |
| Local parsing with no API key                         | [LiteParse](/liteparse/index.md), the open-source parser     |
| Scans, tables, charts, forms, 130+ formats            | Parse, as above                                              |
| Structured fields from documents                      | Extract, as above                                            |
| A managed index that stays in sync with a data source | [Index](/llamaparse/cloud-index-v2/getting_started/index.md) |

## See also

- [Parse getting started](/llamaparse/parse/getting_started/index.md)
- [Framework starter tutorial](/python/framework/getting_started/starter_example/index.md)
- [LlamaIndex Framework overview](/python/framework/index.md)
