Extract
Create Extract Job
List Extract Jobs
Get Extract Job
Delete Extract Job
Cancel Extract Job
Validate Extraction Schema
Generate Extraction Schema
ModelsExpand Collapse
class ExtractConfiguration:
Extract configuration combining parse and extract settings.
Optional<Boolean> citeSources
Include citations in results. Returned under extract_metadata (auto-included when set). Text-level on turbo (no bounding boxes).
Optional<Boolean> confidenceScores
Include confidence scores in results. Returned under extract_metadata (auto-included when set).
Optional<String> parseConfigId
Saved parse configuration ID to control how the document is parsed before extraction. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<ParseTier> parseTier
Parse tier to use before extraction. Defaults to the extract tier if not specified. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<List<String>> sheetNames
Optional worksheet names to extract when spreadsheet_mode is on. Overrides target_pages for spreadsheets; omit to extract every sheet. Names are matched exactly (case-sensitive) — pass them as a list, e.g. [“Sheet 1”, “My Sheet”].
Optional<Boolean> spreadsheetMode
Beta. When true, extract structured data directly from a spreadsheet workbook (.xlsx/.xls/.csv) — the agent reads cells straight from the workbook instead of the standard document path. Off by default (spreadsheets keep the standard path). Requires the agentic_plus tier. Billed on the standard per-page extract rate, against a page count derived from workbook size. Citations and confidence scores are not available in this mode.
Optional<String> targetPages
Comma-separated page numbers or ranges to process (1-based). Omit to process all pages.
Optional<String> version
Use ‘latest’ for the latest release for the selected tier or a date string (YYYY-MM-DD format) to pin to the nearest release at or before that date. Job responses always report the concrete resolved version the job runs, fixed at job creation; saved configurations keep the value as provided.
class ExtractJobMetadata:
Extraction metadata.
Metadata for extracted fields including document, page, and row level info.
Optional<DocumentMetadata> documentMetadata
Per-field metadata keyed by field name from your schema. Scalar fields (e.g. vendor) map to a FieldMetadataEntry with citation and confidence. Array fields (e.g. items) map to a list where each element contains per-sub-field FieldMetadataEntry objects, indexed by array position. Nested objects contain sub-field entries recursively.
class ExtractV2Job:
An extraction job.
String status
Current job status.
PENDING— queued, not yet startedRUNNING— actively processingCOMPLETED— finished successfullyFAILED— terminated with an errorCANCELLED— cancelled by user
Extract configuration combining parse and extract settings.
Optional<Boolean> citeSources
Include citations in results. Returned under extract_metadata (auto-included when set). Text-level on turbo (no bounding boxes).
Optional<Boolean> confidenceScores
Include confidence scores in results. Returned under extract_metadata (auto-included when set).
Optional<String> parseConfigId
Saved parse configuration ID to control how the document is parsed before extraction. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<ParseTier> parseTier
Parse tier to use before extraction. Defaults to the extract tier if not specified. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<List<String>> sheetNames
Optional worksheet names to extract when spreadsheet_mode is on. Overrides target_pages for spreadsheets; omit to extract every sheet. Names are matched exactly (case-sensitive) — pass them as a list, e.g. [“Sheet 1”, “My Sheet”].
Optional<Boolean> spreadsheetMode
Beta. When true, extract structured data directly from a spreadsheet workbook (.xlsx/.xls/.csv) — the agent reads cells straight from the workbook instead of the standard document path. Off by default (spreadsheets keep the standard path). Requires the agentic_plus tier. Billed on the standard per-page extract rate, against a page count derived from workbook size. Citations and confidence scores are not available in this mode.
Optional<String> targetPages
Comma-separated page numbers or ranges to process (1-based). Omit to process all pages.
Optional<String> version
Use ‘latest’ for the latest release for the selected tier or a date string (YYYY-MM-DD format) to pin to the nearest release at or before that date. Job responses always report the concrete resolved version the job runs, fixed at job creation; saved configurations keep the value as provided.
Extraction metadata.
Metadata for extracted fields including document, page, and row level info.
Optional<DocumentMetadata> documentMetadata
Per-field metadata keyed by field name from your schema. Scalar fields (e.g. vendor) map to a FieldMetadataEntry with citation and confidence. Array fields (e.g. items) map to a list where each element contains per-sub-field FieldMetadataEntry objects, indexed by array position. Nested objects contain sub-field entries recursively.
Optional<ExtractResult> extractResult
Extracted data conforming to the data_schema. Returns a single object for per_doc, or an array for per_page / per_table_row.
class ExtractV2JobCreate:
Request to create an extraction job. Provide configuration_id or inline configuration.
Extract configuration combining parse and extract settings.
Optional<Boolean> citeSources
Include citations in results. Returned under extract_metadata (auto-included when set). Text-level on turbo (no bounding boxes).
Optional<Boolean> confidenceScores
Include confidence scores in results. Returned under extract_metadata (auto-included when set).
Optional<String> parseConfigId
Saved parse configuration ID to control how the document is parsed before extraction. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<ParseTier> parseTier
Parse tier to use before extraction. Defaults to the extract tier if not specified. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<List<String>> sheetNames
Optional worksheet names to extract when spreadsheet_mode is on. Overrides target_pages for spreadsheets; omit to extract every sheet. Names are matched exactly (case-sensitive) — pass them as a list, e.g. [“Sheet 1”, “My Sheet”].
Optional<Boolean> spreadsheetMode
Beta. When true, extract structured data directly from a spreadsheet workbook (.xlsx/.xls/.csv) — the agent reads cells straight from the workbook instead of the standard document path. Off by default (spreadsheets keep the standard path). Requires the agentic_plus tier. Billed on the standard per-page extract rate, against a page count derived from workbook size. Citations and confidence scores are not available in this mode.
Optional<String> targetPages
Comma-separated page numbers or ranges to process (1-based). Omit to process all pages.
Optional<String> version
Use ‘latest’ for the latest release for the selected tier or a date string (YYYY-MM-DD format) to pin to the nearest release at or before that date. Job responses always report the concrete resolved version the job runs, fixed at job creation; saved configurations keep the value as provided.
Optional<List<String>> webhookConfigurationIds
IDs of saved webhook configurations to notify for this job.
Optional<List<WebhookConfiguration>> webhookConfigurations
Outbound webhook endpoints to notify on job status changes
Optional<List<WebhookEvent>> webhookEvents
Events to subscribe to (e.g. ‘parse.success’, ‘extract.error’). If null, all events are delivered.
Optional<WebhookHeaders> webhookHeaders
Custom HTTP headers sent with each webhook request (e.g. auth tokens)
Optional<String> webhookOutputFormat
Response format sent to the webhook: ‘string’ (default) or ‘json’
Optional<String> webhookSigningSecret
Shared signing secret used to sign webhook deliveries. When set, each request includes an HMAC-SHA256 signature of the request body in the ‘LC-Signature’ header (value ‘sha256=
class ExtractV2JobQueryResponse:
Paginated list of extraction jobs.
List<ExtractV2Job> items
The list of items.
String status
Current job status.
PENDING— queued, not yet startedRUNNING— actively processingCOMPLETED— finished successfullyFAILED— terminated with an errorCANCELLED— cancelled by user
Extract configuration combining parse and extract settings.
Optional<Boolean> citeSources
Include citations in results. Returned under extract_metadata (auto-included when set). Text-level on turbo (no bounding boxes).
Optional<Boolean> confidenceScores
Include confidence scores in results. Returned under extract_metadata (auto-included when set).
Optional<String> parseConfigId
Saved parse configuration ID to control how the document is parsed before extraction. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<ParseTier> parseTier
Parse tier to use before extraction. Defaults to the extract tier if not specified. Turbo extract does not support parse configuration or produce a parse output; use another tier if your workflow requires parsed text.
Optional<List<String>> sheetNames
Optional worksheet names to extract when spreadsheet_mode is on. Overrides target_pages for spreadsheets; omit to extract every sheet. Names are matched exactly (case-sensitive) — pass them as a list, e.g. [“Sheet 1”, “My Sheet”].
Optional<Boolean> spreadsheetMode
Beta. When true, extract structured data directly from a spreadsheet workbook (.xlsx/.xls/.csv) — the agent reads cells straight from the workbook instead of the standard document path. Off by default (spreadsheets keep the standard path). Requires the agentic_plus tier. Billed on the standard per-page extract rate, against a page count derived from workbook size. Citations and confidence scores are not available in this mode.
Optional<String> targetPages
Comma-separated page numbers or ranges to process (1-based). Omit to process all pages.
Optional<String> version
Use ‘latest’ for the latest release for the selected tier or a date string (YYYY-MM-DD format) to pin to the nearest release at or before that date. Job responses always report the concrete resolved version the job runs, fixed at job creation; saved configurations keep the value as provided.
Extraction metadata.
Metadata for extracted fields including document, page, and row level info.
Optional<DocumentMetadata> documentMetadata
Per-field metadata keyed by field name from your schema. Scalar fields (e.g. vendor) map to a FieldMetadataEntry with citation and confidence. Array fields (e.g. items) map to a list where each element contains per-sub-field FieldMetadataEntry objects, indexed by array position. Nested objects contain sub-field entries recursively.
Optional<ExtractResult> extractResult
Extracted data conforming to the data_schema. Returns a single object for per_doc, or an array for per_page / per_table_row.
class ExtractedFieldMetadata:
Metadata for extracted fields including document, page, and row level info.
Optional<DocumentMetadata> documentMetadata
Per-field metadata keyed by field name from your schema. Scalar fields (e.g. vendor) map to a FieldMetadataEntry with citation and confidence. Array fields (e.g. items) map to a list where each element contains per-sub-field FieldMetadataEntry objects, indexed by array position. Nested objects contain sub-field entries recursively.