Menu
Guides

Deleting documents in bulk in Master Data

Learn how to delete all Master Data documents matching a filter in a single asynchronous job, and how to follow that job until it finishes.

8 min read

In this guide, you will learn how to delete every document that matches a filter in a Master Data data entity, using a single asynchronous deletion job. The job is available for both Master Data v1 and Master Data v2 data entities. When you need to remove everything matching one filter, use this job instead of the scroll and per-document DELETE loop described in Deleting documents in Master Data v1.

This feature is in open beta.

This guide focuses on the create, poll, and confirm flow. For the full operation set, parameters, schemas, and errors, see the API reference in the following table:

OperationMaster Data v1Master Data v2
Create bulk document deletion jobv1v2
Get bulk document deletion job statusv1v2

This action cannot be undone. Deleted documents are permanently lost. Confirm how many documents your filter matches before you create the job, so you can compare that number with the DocumentsDeleted value at the end.

To erase data for one specific customer for privacy reasons, follow Erasing customer data, which uses search and per-document deletion. This guide does not replace that flow.

Before you begin

  • Both the create and status requests require a valid user token in the VtexIdclientAutCookie header.
  • The filter field is required and uses the same syntax as the data entity search filter. Wildcards aren't allowed.
  • Every field in the filter must be indexed. Internal fields such as createdIn are already indexed. Custom fields follow a different rule in each Master Data version:
    • In Master Data v1, a custom field is indexed when it exists in the data entity and has isSearchable enabled. You don't need to send the schema field.
    • In Master Data v2, a custom field is indexed when it is listed in the v-indexed array of a schema, and the request must declare that schema in the schema field.
  • There is no limit to how many documents a single job can delete, but deleting hundreds of thousands of documents or more at once raises the chance of failure. Whenever possible, narrow the filter down to batches of a few tens of thousands of documents. For example, in a data entity with 10 million documents, run several jobs partitioned by a date field instead of a single filter that matches everything.
  • Only one deletion job can be active per account and data entity at a time. While a job is InProgress, creating another job for the same data entity returns 409. Follow the existing job as described in step 2 and send the new request once that job reaches Success or Failed. If a job appears stuck, Master Data automatically allows a new job for that data entity after 12 hours.
  • Before you create the job, search the data entity with the same filter criteria, and the same schema in the case of Master Data v2, record the number of matching documents as your baseline count, and save a few of the returned document IDs. You'll use these in step 3. To learn the query patterns and count how many documents a filter matches, see Extracting data from Master Data with search and scroll.

How it works

Deletion is asynchronous and runs in three steps. The create request deletes nothing by itself. It only creates the job and returns a JobId.

  1. Create the job with Create bulk document deletion job.
  2. Poll Get bulk document deletion job status until the job reaches Success or Failed.
  3. Confirm the result against your baseline count and a follow-up search.

Instructions

Step 1 - Create the deletion job

Send a POST request to Create bulk document deletion job, or to the Master Data v1 equivalent.

The following request body filters on an internal field:


_10
{
_10
"filter": "createdIn > 2026-08-16"
_10
}

In Master Data v2, filtering on a custom indexed field also requires the schema that declares the field as indexed:


_10
{
_10
"filter": "number = 1337",
_10
"schema": "indexed-fields"
_10
}

In Master Data v1, filtering on a custom field doesn't require schema. The field only needs isSearchable enabled:


_10
{
_10
"filter": "isActive = true"
_10
}

A successful request returns HTTP status 202 Accepted:


_10
{
_10
"JobId": "01M08C83Z5D91S3CA0A15SRNV8"
_10
}

Save the JobId. It's the only way to track the operation.

Step 2 - Follow the job status

Send a GET request to Get bulk document deletion job status, or to the Master Data v1 equivalent, using the JobId from step 1.

Poll until Status is Success or Failed. Example response for a finished job:


_10
{
_10
"JobId": "01KZEEJ9ZRQ4XMMMBXNEX70EF4",
_10
"Entity": "bulk_delete_test",
_10
"Status": "Success",
_10
"Filter": "number < 100",
_10
"DocumentsDeleted": 19,
_10
"CreatedAt": "2026-08-07T15:48:15.4804028Z",
_10
"UpdatedAt": "2026-08-12T19:31:10.654298Z"
_10
}

Each status requires a different action:

  • InProgress: The job is running. Keep polling. Don't submit another job for this data entity.
  • Success: The job deleted all matching documents. Compare DocumentsDeleted with your baseline count, as described in step 3.
  • Failed: Processing stopped after the job exhausted its internal retries. Deletion is partial. Search the data entity with the same filter to see which documents remain, then create a new job with that filter to finish the deletion. For errors that reject a new job, see Create bulk document deletion job or the Master Data v1 equivalent.

Failed doesn't mean that nothing was deleted. Documents from batches the job already completed are permanently gone.

Step 3 - Confirm the deletion result

  1. Read DocumentsDeleted from the job status response and compare it with the baseline count you recorded before creating the job.
  2. Search the data entity again with the same filter criteria, and the same schema in the case of Master Data v2, using Search documents or the Master Data v1 equivalent. It should return no documents.
  3. Send a GET request for one of the document IDs you saved, using Get document or the Master Data v1 equivalent. It should return an empty response.

After the documents are deleted, they are no longer counted in stored volume.

Error reference

POST /api/dataentities/{name}/delete

Master Data v1 and Master Data v2 validate the request on separate code paths, so some 400 errors are specific to one version, as indicated in the Version column.

HTTP statusVersionMessage or causeAction
400v1 and v2The {filter} field is requiredSend the filter field in the body.
400v1 and v2Wildcard filters are not allowed for bulk delete.Rewrite the filter using exact or range conditions over indexed fields.
400v2 onlyThe field {field} is not an internal indexed field, so the 'schema' that declares it as indexed must be providedAdd the schema that declares the field as indexed to the body, as shown in step 1.
400v2 onlyThe field '{field}' is not indexed in the schema '{schema}'...Mark the field as indexed in the schema and wait for reindexing, or filter on a field that is already indexed.
400v1 onlyThe field {field} does not exist for the data entity '{entity}'Correct the field name in the filter.
400v1 onlyThe field {field} of the data entity {entity} is not indexed (isSearchable is not enabled)...Enable isSearchable on the field and wait for reindexing, or use another indexed field.
400v1 and v2Invalid data entity name.Correct the data entity name in the URL. Invalid characters are rejected.
400v1 and v2The request body is empty or isn't valid JSON, and the Content-Type header is present. This response follows the standard validation format and has no Message field.Send a valid JSON body containing the filter field.
409v1 and v2Bulk deletion job creation failed for entity {entity}. A job is already InProgress for this account and data entity.Follow the existing job as described in step 2 and send the request again once it reaches a terminal state. If a job appears stuck, the lock is released automatically after 12 hours.
415v1 and v2The Content-Type header is missing, or the media type isn't supported.Send the request with the Content-Type: application/json header.
5xxv1 and v2Transient infrastructure failure during job creation.No residual state is left behind. Send the request again.

GET /api/dataentities/{name}/delete/jobs/{jobId}

HTTP statusCauseAction
404The jobId does not exist for this account and data entity, or the job was created more than 60 days ago. Job records are kept for 60 days after creation.Check the jobId and the data entity name in the URL, and make sure the an query parameter points to the correct account.

Next steps