## Create Split Job

`split.create(SplitCreateParams**kwargs)  -> SplitCreateResponse`

**post** `/api/v1/split/jobs`

Create a document split job.

### Parameters

- `file_input: str`

  File ID or parse job ID

- `organization_id: Optional[str]`

- `project_id: Optional[str]`

- `configuration: Optional[Configuration]`

  Split configuration with categories and splitting strategy.

  - `categories: Iterable[SplitCategoryParam]`

    Categories to split documents into.

    - `name: str`

      Name of the category.

    - `description: Optional[str]`

      Optional description of what content belongs in this category.

  - `splitting_strategy: Optional[ConfigurationSplittingStrategy]`

    Strategy for splitting documents.

    - `allow_uncategorized: Optional[Literal["forbid", "include", "omit"]]`

      Controls handling of pages that don't match any category. 'include': pages can be grouped as 'uncategorized' and included in results. 'forbid': all pages must be assigned to a defined category. 'omit': pages can be classified as 'uncategorized' but are excluded from results.

      - `"forbid"`

      - `"include"`

      - `"omit"`

    - `custom_instructions: Optional[str]`

      Free-form guidance for where segment boundaries are placed.

    - `min_pages_per_split: Optional[int]`

      Minimum pages per segment. Shorter segments are merged into an adjacent segment; 1 disables merging.

- `configuration_id: Optional[str]`

  Saved configuration ID

- `transaction_id: Optional[str]`

  Idempotency key scoped to the project. Reusing a key returns the original job; the new request body is ignored.

- `webhook_configuration_ids: Optional[Sequence[str]]`

  IDs of saved webhook configurations to notify for this job.

- `webhook_configurations: Optional[Iterable[WebhookConfiguration]]`

  Outbound webhook endpoints to notify on job status changes

  - `webhook_events: Optional[List[Literal["batch.cancelled", "batch.error", "batch.pending", 30 more]]]`

    Events to subscribe to (e.g. 'parse.success', 'extract.error'). If null, all events are delivered.

    - `"batch.cancelled"`

    - `"batch.error"`

    - `"batch.pending"`

    - `"batch.running"`

    - `"batch.success"`

    - `"classify.cancelled"`

    - `"classify.error"`

    - `"classify.partial_success"`

    - `"classify.pending"`

    - `"classify.running"`

    - `"classify.success"`

    - `"extract.cancelled"`

    - `"extract.error"`

    - `"extract.partial_success"`

    - `"extract.pending"`

    - `"extract.success"`

    - `"parse.cancelled"`

    - `"parse.error"`

    - `"parse.partial_success"`

    - `"parse.pending"`

    - `"parse.running"`

    - `"parse.success"`

    - `"sheets.cancelled"`

    - `"sheets.error"`

    - `"sheets.partial_success"`

    - `"sheets.pending"`

    - `"sheets.success"`

    - `"split.cancelled"`

    - `"split.error"`

    - `"split.pending"`

    - `"split.processing"`

    - `"split.success"`

    - `"unmapped_event"`

  - `webhook_headers: Optional[Dict[str, str]]`

    Custom HTTP headers sent with each webhook request (e.g. auth tokens)

  - `webhook_output_format: Optional[str]`

    Response format sent to the webhook: 'string' (default) or 'json'

  - `webhook_signing_secret: Optional[str]`

    Shared signing secret used to sign webhook deliveries. When set, each request includes an HMAC-SHA256 signature of the request body in the 'LC-Signature' header (value 'sha256=<hex>'). Recompute the HMAC over the raw request body with this secret to verify the delivery is authentic.

  - `webhook_url: Optional[str]`

    URL to receive webhook POST notifications

### Returns

- `class SplitCreateResponse: …`

  A split job.

  - `id: str`

    Unique identifier for the split job.

  - `categories: List[SplitCategory]`

    Categories used for splitting.

    - `name: str`

      Name of the category.

    - `description: Optional[str]`

      Optional description of what content belongs in this category.

  - `document_input_type: Literal["file_id", "parse_job_id", "url"]`

    Whether the input was a file or parse job

    - `"file_id"`

    - `"parse_job_id"`

    - `"url"`

  - `file_input: str`

    File ID or parse job ID

  - `project_id: str`

    Project this job belongs to.

  - `status: str`

    Current job status. Valid values are: pending, processing, completed, failed, cancelled.

  - `user_id: str`

    User who created this job.

  - `configuration_id: Optional[str]`

    Split configuration ID used for this job.

  - `created_at: Optional[datetime]`

    Creation datetime

  - `error_message: Optional[str]`

    Error message if the job failed.

  - `result: Optional[SplitResultResponse]`

    Result of a completed split job.

    - `segments: List[SplitSegmentResponse]`

      List of document segments.

      - `category: str`

        Category name this split belongs to.

      - `confidence_category: str`

        Categorical confidence level. Valid values are: high, medium, low.

      - `pages: List[int]`

        1-indexed page numbers in this split.

  - `splitting_strategy: Optional[SplittingStrategy]`

    Strategy used for splitting.

    - `allow_uncategorized: Optional[Literal["forbid", "include", "omit"]]`

      Controls handling of pages that don't match any category. 'include': pages can be grouped as 'uncategorized' and included in results. 'forbid': all pages must be assigned to a defined category. 'omit': pages can be classified as 'uncategorized' but are excluded from results.

      - `"forbid"`

      - `"include"`

      - `"omit"`

    - `custom_instructions: Optional[str]`

      Free-form guidance for where segment boundaries are placed.

    - `min_pages_per_split: Optional[int]`

      Minimum pages per segment. Shorter segments are merged into an adjacent segment; 1 disables merging.

  - `transaction_id: Optional[str]`

    Idempotency key scoped to the project, if one was provided.

  - `updated_at: Optional[datetime]`

    Update datetime

### Example

```python
import os
from llama_cloud import LlamaCloud

client = LlamaCloud(
    api_key=os.environ.get("LLAMA_CLOUD_API_KEY"),  # This is the default and can be omitted
)
split = client.split.create(
    file_input="dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
)
print(split.id)
```

#### Response

```json
{
  "id": "id",
  "categories": [
    {
      "name": "x",
      "description": "x"
    }
  ],
  "document_input_type": "file_id",
  "file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
  "project_id": "project_id",
  "status": "status",
  "user_id": "user_id",
  "configuration_id": "configuration_id",
  "created_at": "2019-12-27T18:11:19.117Z",
  "error_message": "error_message",
  "result": {
    "segments": [
      {
        "category": "category",
        "confidence_category": "confidence_category",
        "pages": [
          0
        ]
      }
    ]
  },
  "splitting_strategy": {
    "allow_uncategorized": "forbid",
    "custom_instructions": "Start a new segment at every signature page.",
    "min_pages_per_split": 1
  },
  "transaction_id": "transaction_id",
  "updated_at": "2019-12-27T18:11:19.117Z"
}
```
