01 / 08

How would you handle a 'Bulk Write' failure mid-way through 10,000 documents?

When a bulk write operation fails midway through 10,000 documents, your response strategy depends entirely on whether the operation was ordered or unordered. Ordered operations stop at the first error, potentially leaving a partially written batch. Unordered operations continue processing remaining documents even after errors, providing a partial result. MongoDB drivers provide specific exception types like MongoBulkWriteException or BulkWriteCommandException that contain detailed error information and partial results, enabling you to identify exactly which documents failed and why .

By default, bulkWrite() performs ordered operations, executing documents serially in the order provided. If an error occurs during an ordered bulk write, MongoDB returns without processing any remaining operations. For unordered operations (set ordered: false), MongoDB can execute operations in parallel and continues processing remaining operations even if some fail. This fundamental difference determines how you handle failures—ordered operations require restarting from the failure point, while unordered operations give you a partial result with failed operations identified by index.

Handling Bulk Write Exceptions

For ordered operations, the first error stops execution, leaving a partial write. The simplest recovery approach is to split your 10,000 documents into smaller batches (e.g., 500-1000 documents per batch). This way, if a batch fails, only that batch needs retry logic, and you can track success across batches. After identifying the failing document, you can isolate and fix the issue (like a duplicate key violation or validation error), then retry the remainder of the batch starting from that point.

With unordered operations, MongoDB continues processing even after errors. The driver throws a bulk write exception that contains a partial result object. This partial result includes counts of successfully processed operations (insertedCount, modifiedCount, deletedCount) and an array of write errors with the index of each failed operation. Your recovery strategy should: capture the failed operation indices from the exception, extract the corresponding documents from your original array, fix any data issues (like duplicate keys or schema violations), and retry only those failed operations in a new bulk write.

If you need true atomicity—either all 10,000 documents succeed or none do—you cannot rely on bulk write alone. Bulk write operations are not atomic across multiple documents; each individual write is atomic, but the batch as a whole is not. For all-or-nothing semantics across the entire batch, you must use multi-document transactions. However, transactions have performance overhead and are generally not recommended for batches this large. The practical approach is to design for idempotency and implement retry logic with partial results [citation:9].

Preventive Strategies
  1. 1

    Batch splitting: Break 10,000 documents into smaller batches (500-1000) to limit the impact of failures and simplify retry logic [citation:3].

  2. 2

    Pre-validation: Validate documents against schema requirements and check for duplicate keys before executing the bulk write to reduce failure rates.

  3. 3

    Idempotent operations: Design operations to be idempotent, allowing safe retries without causing duplicate data.

  4. 4

    Index analysis: Ensure unique indexes are correctly configured and that your data won't violate them, as duplicate key errors (code 11000) are the most common bulk write failures [citation:3][citation:10].

Batch Splitting Implementation
Difficulty: 8/10
Topics: bulk write atomicity, error handling in batch operations, retry strategies with partial failure

Scenario Questions

0-2 years experience
  1. 1

    If you're inserting 10,000 documents and the 5,327th one fails due to a duplicate key, what does MongoDB do by default, and how would you find out which one failed?

  2. 2

    You're writing a script to bulk insert user data and it crashes halfway — what’s the simplest way to resume without re-inserting everything you already did?

2-5 years experience
  1. 1

    Our user onboarding pipeline bulk-inserts profile data, but sometimes 200 out of 10,000 docs fail due to validation errors — users are complaining their accounts are missing. How would you debug and fix this?

  2. 2

    A batch job failed mid-way during a nightly data sync, and now support is flooded with tickets about missing records. How do you determine what succeeded, what failed, and how to recover without double-inserting?

5-8 years experience
  1. 1

    You're designing a high-throughput ingestion system that processes 50K documents per batch — how do you balance throughput, error isolation, and recovery complexity when bulk writes can fail partially?

  2. 2

    In a distributed system, bulk writes sometimes fail due to network timeouts or transient MongoDB errors. How would you design a retry mechanism that avoids duplicates, preserves order when needed, and doesn’t overwhelm the database?

8+ years experience
  1. 1

    We’re migrating from a legacy system that assumes all-or-nothing bulk writes, but MongoDB’s partial failure model breaks our assumptions — how would you redesign the ingestion pipeline across teams without causing data inconsistency or downtime?

  2. 2

    Our bulk write pipeline handles millions of documents daily, and partial failures are now a recurring operational burden. How would you architect a self-healing, auditable system that minimizes manual intervention and ensures eventual consistency across downstream consumers?

Follow-up Questions

  • How would you verify that no documents were lost or duplicated after a failure?
  • What metrics would you track to detect recurring bulk write failures in production?
  • Would you change your approach if the 10,000 documents came from user uploads vs. system-generated data?