Apache Iceberg makes data lake storage update-able. However, this gets tricky when you’re writing change-data-capture (CDC) or streaming ingestion jobs. You might have run into:
Compaction jobs failing with Cannot commit, found new delete for replaced data file
Ingestion commits being rejected with Missing required files to delete
Degraded query performance, often due to compaction jobs not committing
To understand why this happens, we need to look at how Iceberg handles compaction and streaming updates, and the trade-offs involved. Let's start with what we're trying to achieve.
Fresh, fast-to-query, interoperable - choose two?
In streaming use cases - and especially for production scenarios such as operational dashboards - we typically want three things:
Low-latency: tables should reflect the business reality within seconds or minutes of a change happening
Fast queries: the underlying files should be compacted regularly so queries stay fast and cheap, rather than degrading as small files and delete files pile up.
Interoperable: tables should be readable in your analytics tools of choice such as Snowflake, Databricks, Redshift, etc.
Doing two of the three is easy enough in Iceberg. Achieving all three is surprisingly tricky. One of the main complications is dealing with clashes between compaction jobs and upserts, which lead to failed commits and degraded performance. Workarounds typically involve scheduling - i.e., running maintenance off hours or widening the retry window - and placing both operations under a single writer. However, as we will explain below, these do not actually solve the issue, and end up costing you on one of the three dimensions.
How Updates Go Wrong: Position Deletes, Equality Deletes, Appends
Let's look at a concrete example. Suppose you’re an online store, running a pipeline that replicates an operational database into Iceberg. A user removes an item from their cart, and you want the Iceberg table to reflect that within seconds. How does that update happen?
Iceberg can't edit a row in place. Object stores like S3 don't let you modify part of an object - you can only write a new one. Iceberg's snapshot model also depends on immutability: snapshots reference data files by path, while snapshot isolation and time travel depend on older files never changing. An update has to be expressed as either a delete plus an insert, or a full rewrite of the file (copy-on-write).
The delete is a separate file telling the query engine which rows to ignore - either by position (row X in file Y) or by equality (where primary_key = x), both of which are supported by the Iceberg spec. The delete takes effect during read time, as the query engine is instructed to ignore deleted rows. The record is then purged from the table’s underlying data during compaction - the deleted rows are removed from the compacted file, and the new file is swapped for the old files at the metadata level.

Conflicts happen when a delete references a file that a compaction job is rewriting. In a streaming pipeline, this often happens because a compaction job might take up to a few minutes to run, while new delete files arrive continuously. Iceberg's rule at commit time is that a compaction must never lose a delete that has already been committed. If swapping in the rewritten file would bring deleted rows back, Iceberg fails the compaction commit. (If your pipeline is append-only, none of this applies: new data files don't reference existing ones, so they never conflict with compaction. But for CDC use cases, where rows get updated and deleted, append-only isn't an option.)
Whether a conflict gets triggered - and what it costs to avoid triggering it - depends on which kind of delete you're using:
Compaction vs. Position Deletes / Deletion Vectors: No Easy Solution
Position-based deletes (deletion vectors in Iceberg v3, position deletes in v2) instruct the query engine to ignore certain rows in certain files. A common way to write position deletes for streaming data is a Spark MERGE INTO operation. Changes stream into a staging table, and the merge applies them to the target table as often as possible, typically every few minutes. (Writing position deletes directly from a streaming job is hard, since finding the file and row of every updated record means a lookup against the table on every write.)
Periodic merges are a data warehouse technique that adds latency compared to true streaming ingestion, since your table is only as fresh as the last merge. However, the main problem is compaction conflicts.
To keep things simple, let's say we compact every three files that land on S3, and the compaction job takes 60 seconds to run. 30 seconds in, a MERGE INTO deletes row 0 in file 1 and row 1 in file 3. The delete is valid and committed, but the compaction job, which is still running, does not account for it.

After another 30 seconds, compaction is ready to commit a new file 4 (containing the rows from files 1, 2, and 3, as they were before the delete). But first, the engine writing the compaction job checks what has been committed since compaction started, and finds a delete file that references files 1 and 3. If committed, file 4 would contain the two deleted rows and the delete file would point at nothing. So it refuses: the commit fails with Cannot commit, found new delete for replaced data file. (If you’ve enabled use-starting-sequence-number, which we cover below, the exact error message might be Cannot commit, found new position delete for replaced data file.)
It can also happen the other way round. If compaction commits first, the merge fails instead, because the files it planned against are no longer in the table. You'll see Cannot commit, missing data files, or Missing required files to delete: s3a://…/data-00001.parquet if the table uses copy-on-write. Either way, one of the two jobs has to run again.
Retrying doesn't fix it. You can configure your compactor to re-plan failed jobs around the latest snapshot, which now includes the delete. But the same conflict is very likely to happen again, because new row updates have been committed to the table since. It's actually more likely, because the next compaction attempt covers a larger set of files and takes longer to run. And while you’re waiting for a successful retry, queries run against un-compacted data, which degrades performance.
Neither does scheduling. You can compact off-hours, pause ingestion while compaction runs, or shorten compaction jobs to narrow the window. These reduce failed commits, but at a price: you're either adding latency (because ingestion waits until compaction finishes) or increasing compute costs (by running more, smaller compactions). If you're running a periodic merge, the squeeze comes from both ends: merging more often for fresher data leaves a smaller window for compaction, while handing it more delete files to resolve.
None of these is airtight either. Scheduling logic needs maintaining, and collisions still happen. Running the merge and the compaction from the same scheduler or process can avoid collisions, but only by making them run in sequence, which re-introduces latency costs.
Compaction vs. Equality Deletes: Sequencing Works, but Performance Suffers
If you’re writing deletes using a streaming pipeline - e.g., writing CDC into Iceberg with Flink - you’re most likely using equality deletes. In an equality delete, you’re using a where predicate instead of a file location. The delete file tells the compaction process to ignore certain rows based on record values. In CDC pipelines, this is almost always the primary key from the source database, e.g. DELETE WHERE id = 32412. Equality deletes are cheap to write because the writer only needs the key and doesn't have to read anything from the table. That's why they're the natural fit for streaming.
When reading the data, the compaction or query engine still needs to know which files this delete applies to, and this is done through sequencing: every commit is assigned a sequence number, which is recorded in the table's snapshot metadata, and every data and delete file written in that commit inherits it. An equality delete only applies to data files with a lower sequence number than its own - otherwise future versions of the same row id would continue to get targeted by the delete, which would remove rows we intended to keep.
When files land mid-compaction, sequencing can go sideways. Returning to our previous example, the first three files would be sequenced 1, 2, 3; the equality delete file (or any other files that landed during the compaction job) would be sequenced 4; the compacted files, which lands after the delete, would be 5.

Without any fix, this compaction would fail for the same reason as the position-delete case: if file 4 were committed with sequence number 5, the delete at 4 would stop applying and the deleted rows would come back from the dead - so instead, the compaction process fails to commit. However:
Iceberg has a fix on the write side. Enabling use-starting-sequence-number tells Iceberg to write the compacted file with the sequence number of the snapshot it started from - in our case 3 - and to record that number explicitly in the file's manifest entry, rather than letting it inherit the compaction commit's own number (5). The delete at 4 now applies to the compacted file, and the commit goes through. This is set to true by default if you're using Spark's rewrite_data_files for compaction, but must be explicitly enabled if using Flink via useStartingSequenceNumber.

This does in fact solve the compaction conflict. However, the trade-off is on the read side: that equality deletes degrade performance to the extent that major query engines do not support them . An equality delete is a predicate that needs to be evaluated: on every scan, the engine loads the ID lists from the relevant delete files into memory and checks each row against them. This is computationally expensive and hurts query performance, to the extent that most major query engines have decided it isn't worth supporting. For example, Snowflake will not read an Iceberg table that contains equality delete files; the Iceberg community also voted to forbid writing new equality deletes in v4 tables. In other words, you're paying a significant cost when it comes to the actual usefulness of your data.
To Summarize So Far...
When using position deletes in low latency streaming, compaction conflicts are not a timing issue. Scheduling and managing the compaction window can reduce conflicts, but the cost will be worse performance or higher latency.
Equality deletes can cause compaction conflicts due to sequencing. Enabling use-starting-sequence-number will typically prevent these conflicts.
However, equality deletes degrade performance and are not supported by major query engines, including Snowflake.
Hence: compaction conflicts in streaming ingestion are a design issue. Every workaround that makes collisions rarer costs you either latency or query performance. What does an actually-working solution look like?
Emerging Patterns for Solving the Design Problem
Since the problem is caused by a conflict between incompatible operations, not their timing, that's where the fix needs to apply. We also can't rely on equality deletes, which are slow to query and unsupported by major query engines.
What we can do is design the write path so that collisions can't happen - so that a position delete and a compaction of the file it points at can never run against the same files at the same time. A single writer is not a sufficient fix, as one process can still issue a delete and a compaction that collide, and avoiding these collisions means adding more latency. The operations should be able to run without collisions, regardless of their timing.
This is the approach we've taken at Etleap: we stream CDC updates through Flink into Iceberg tables that Snowflake reads directly, with end-to-end latency of around two minutes, and with no possibility of conflicts between ingestion and compaction.
Without getting into the weeds, we use Iceberg branching to isolate ingestion from the tables your query engines read, along with a few other tricks to keep the two from colliding. It comes with its own pitfalls and takes some careful engineering, and there's a small, fixed amount of work per delete that might get repeated. But that replaces an open-ended cycle of failed compactions and retries with predictable overhead. The Iceberg project itself is also moving in a similar direction (although the proposed solution still relies on locking, which introduces additional latency to the data).
We’ll provide more details about our solution in a future article, and you can also learn more by watching Archie's talk from Iceberg Summit 2026.
Final Thoughts
We hope that this piece has helped you understand the challenges involved in streaming updates, and why conflicts between maintenance and compaction can happen. If you're tuning retry counts and maintenance windows, you're working around a built-in architectural constraint in Iceberg, and these workarounds will only go so far. Scheduling changes how often conflicting operations meet; only the design of the write path changes whether they can.
What should you do next?
If you’ve got a running pipeline and you’re encountering compaction-related errors, you can try implementing one of the fixes outlined above. Your tables should still be correct, but you might have issues with latency and costs.
If you’re still designing your pipeline, decide how deletes get written with compaction in mind from the start. Changing it later means rebuilding the write path.
Subscribe below to read our future piece on how we solved this problem in building our streaming solution at Etleap.