Skip to content

Create Split Job

POST/api/v1/split/jobs

Create a document split job.

Query ParametersExpand Collapse
organization_id: optional string
project_id: optional string
Cookie ParametersExpand Collapse
session: optional string
Body ParametersJSONExpand Collapse
file_input: string

File ID or parse job ID

configuration: optional object { categories, splitting_strategy }

Split configuration with categories and splitting strategy.

categories: array of SplitCategory { name, description }

Categories to split documents into.

name: string

Name of the category.

maxLength200
minLength1
description: optional string

Optional description of what content belongs in this category.

maxLength2000
minLength1
splitting_strategy: optional object { allow_uncategorized, custom_instructions, min_pages_per_split }

Strategy for splitting documents.

allow_uncategorized: optional "forbid" or "include" or "omit"

Controls handling of pages that don’t match any category. ‘include’: pages can be grouped as ‘uncategorized’ and included in results. ‘forbid’: all pages must be assigned to a defined category. ‘omit’: pages can be classified as ‘uncategorized’ but are excluded from results.

One of the following:
"forbid"
"include"
"omit"
custom_instructions: optional string

Free-form guidance for where segment boundaries are placed.

maxLength5000
min_pages_per_split: optional number

Minimum pages per segment. Shorter segments are merged into an adjacent segment; 1 disables merging.

minimum1
configuration_id: optional string

Saved configuration ID

transaction_id: optional string

Idempotency key scoped to the project. Reusing a key returns the original job; the new request body is ignored.

maxLength255
webhook_configuration_ids: optional array of string

IDs of saved webhook configurations to notify for this job.

webhook_configurations: optional array of object { webhook_events, webhook_headers, webhook_output_format, 2 more }

Outbound webhook endpoints to notify on job status changes

webhook_events: optional array of "batch.cancelled" or "batch.error" or "batch.pending" or 30 more

Events to subscribe to (e.g. ‘parse.success’, ‘extract.error’). If null, all events are delivered.

One of the following:
"batch.cancelled"
"batch.error"
"batch.pending"
"batch.running"
"batch.success"
"classify.cancelled"
"classify.error"
"classify.partial_success"
"classify.pending"
"classify.running"
"classify.success"
"extract.cancelled"
"extract.error"
"extract.partial_success"
"extract.pending"
"extract.success"
"parse.cancelled"
"parse.error"
"parse.partial_success"
"parse.pending"
"parse.running"
"parse.success"
"sheets.cancelled"
"sheets.error"
"sheets.partial_success"
"sheets.pending"
"sheets.success"
"split.cancelled"
"split.error"
"split.pending"
"split.processing"
"split.success"
"unmapped_event"
webhook_headers: optional map[string]

Custom HTTP headers sent with each webhook request (e.g. auth tokens)

webhook_output_format: optional string

Response format sent to the webhook: ‘string’ (default) or ‘json’

webhook_signing_secret: optional string

Shared signing secret used to sign webhook deliveries. When set, each request includes an HMAC-SHA256 signature of the request body in the ‘LC-Signature’ header (value ‘sha256=’). Recompute the HMAC over the raw request body with this secret to verify the delivery is authentic.

webhook_url: optional string

URL to receive webhook POST notifications

ReturnsExpand Collapse
id: string

Unique identifier for the split job.

categories: array of SplitCategory { name, description }

Categories used for splitting.

name: string

Name of the category.

maxLength200
minLength1
description: optional string

Optional description of what content belongs in this category.

maxLength2000
minLength1
document_input_type: "file_id" or "parse_job_id" or "url"

Whether the input was a file or parse job

One of the following:
"file_id"
"parse_job_id"
"url"
file_input: string

File ID or parse job ID

project_id: string

Project this job belongs to.

status: string

Current job status. Valid values are: pending, processing, completed, failed, cancelled.

user_id: string

User who created this job.

configuration_id: optional string

Split configuration ID used for this job.

created_at: optional string

Creation datetime

formatdate-time
error_message: optional string

Error message if the job failed.

result: optional SplitResultResponse { segments }

Result of a completed split job.

segments: array of SplitSegmentResponse { category, confidence_category, pages }

List of document segments.

category: string

Category name this split belongs to.

confidence_category: string

Categorical confidence level. Valid values are: high, medium, low.

pages: array of number

1-indexed page numbers in this split.

splitting_strategy: optional object { allow_uncategorized, custom_instructions, min_pages_per_split }

Strategy used for splitting.

allow_uncategorized: optional "forbid" or "include" or "omit"

Controls handling of pages that don’t match any category. ‘include’: pages can be grouped as ‘uncategorized’ and included in results. ‘forbid’: all pages must be assigned to a defined category. ‘omit’: pages can be classified as ‘uncategorized’ but are excluded from results.

One of the following:
"forbid"
"include"
"omit"
custom_instructions: optional string

Free-form guidance for where segment boundaries are placed.

maxLength5000
min_pages_per_split: optional number

Minimum pages per segment. Shorter segments are merged into an adjacent segment; 1 disables merging.

minimum1
transaction_id: optional string

Idempotency key scoped to the project, if one was provided.

updated_at: optional string

Update datetime

formatdate-time

Create Split Job

curl https://api.cloud.llamaindex.ai/api/v1/split/jobs \
    -H 'Content-Type: application/json' \
    -H "Authorization: Bearer $LLAMA_CLOUD_API_KEY" \
    -d '{
          "file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee"
        }'
{
  "id": "id",
  "categories": [
    {
      "name": "x",
      "description": "x"
    }
  ],
  "document_input_type": "file_id",
  "file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
  "project_id": "project_id",
  "status": "status",
  "user_id": "user_id",
  "configuration_id": "configuration_id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "error_message": "error_message",
  "result": {
    "segments": [
      {
        "category": "category",
        "confidence_category": "confidence_category",
        "pages": [
          0
        ]
      }
    ]
  },
  "splitting_strategy": {
    "allow_uncategorized": "forbid",
    "custom_instructions": "Start a new segment at every signature page.",
    "min_pages_per_split": 1
  },
  "transaction_id": "transaction_id",
  "updated_at": "2019-12-27T18:11:19.117Z"
}
Returns Examples
{
  "id": "id",
  "categories": [
    {
      "name": "x",
      "description": "x"
    }
  ],
  "document_input_type": "file_id",
  "file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
  "project_id": "project_id",
  "status": "status",
  "user_id": "user_id",
  "configuration_id": "configuration_id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "error_message": "error_message",
  "result": {
    "segments": [
      {
        "category": "category",
        "confidence_category": "confidence_category",
        "pages": [
          0
        ]
      }
    ]
  },
  "splitting_strategy": {
    "allow_uncategorized": "forbid",
    "custom_instructions": "Start a new segment at every signature page.",
    "min_pages_per_split": 1
  },
  "transaction_id": "transaction_id",
  "updated_at": "2019-12-27T18:11:19.117Z"
}
Note for AI agents: this documentation is built for programmatic access. - Overview of all docs: https://developers.llamaindex.ai/llms.txt - Any page is available as raw Markdown by appending index.md to its URL — e.g. https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md - Agent-friendly REST search APIs live under https://developers.llamaindex.ai/api/ — search (BM25 full-text), grep (regex), read (fetch a page), and list (browse the doc tree). See https://developers.llamaindex.ai/llms.txt for parameters. - A hosted documentation MCP server is available at https://developers.llamaindex.ai/mcp. If you support MCP, you can ask the user to install it for browsing these docs directly (an alternative to the REST API). Setup: https://developers.llamaindex.ai/for-agents/mcp/ - Other LlamaIndex tooling for agents — the LlamaParse Platform MCP server, agent skills and plugins, and the n8n node — is mapped at https://developers.llamaindex.ai/for-agents/