Skip to content

fix: don't let S3A committer, delete and read defaults block native Iceberg writes - #6808

Merged
andygrove merged 1 commit into
apache:mainfrom
andygrove:ignore-s3a-cluster-defaults
Oct 10, 2026
Merged

andygrove merged 1 commit into
apache:mainfrom
andygrove:ignore-s3a-cluster-defaults

Conversation

@andygrove

Copy link
Copy Markdown
Member

Which issue does this PR close?

No issue. This follows up on #6441.

Rationale for this change

#6441 makes a native Iceberg write to S3 fall back to Spark when the Hadoop configuration has an fs.s3a.* setting the native writer doesn't support. It already ignores Hadoop's built-in defaults and a few S3A read settings that Spark seeds into every session. Some other settings are also common in S3 deployments, and any one of them makes every native S3 write fall back:

  • Since Spark 4.1, SparkContext sets fs.s3a.committer.magic.enabled=true and fs.s3a.committer.name=magic for every application when spark-hadoop-cloud is on the classpath (SPARK-47618).
  • Many deployments also set fs.s3a.committer.threads, fs.s3a.experimental.input.fadvise (Hadoop's S3A docs recommend random for columnar formats) or fs.s3a.bulk.delete.page.size cluster-wide.

The fallback reason looks like this:

unsupported Hadoop S3A settings: fs.s3a.bulk.delete.page.size, fs.s3a.committer.magic.enabled, fs.s3a.committer.name, fs.s3a.committer.threads, fs.s3a.experimental.input.fadvise

None of these settings affects an Iceberg data-file write:

  • The native writer doesn't go through S3A, so S3A's delete batching doesn't apply.
  • fadvise is a read hint.
  • Iceberg commits through table metadata, not through a Hadoop output committer.

What changes are included in this PR?

The five keys are added to IgnoredHadoopS3Keys in CometIcebergNativeWrite, alongside the Spark-seeded read settings, with a comment explaining why.

How are these changes tested?

There's a new test in CometIcebergWriteDetectionSuite. It checks that these settings produce no unsupported keys, and that a setting the native writer can't honor (fs.s3a.encryption.algorithm) is still reported next to them. The whole suite passes locally.

…ceberg writes

apache#6441 makes a native Iceberg write to S3 fall back when the Hadoop
configuration has an fs.s3a.* setting the native writer doesn't support.
Spark 4.1 sets fs.s3a.committer.magic.enabled and fs.s3a.committer.name
for every application when spark-hadoop-cloud is on the classpath
(SPARK-47618), and many deployments set fs.s3a.experimental.input.fadvise,
fs.s3a.bulk.delete.page.size or fs.s3a.committer.threads cluster-wide. Any
of them makes every native S3 write fall back.

None of these settings affects an Iceberg data-file write: the native
writer doesn't go through S3A, fadvise is a read hint, and Iceberg commits
through table metadata rather than a Hadoop output committer. Ignore them
like the Spark-seeded read settings already in IgnoredHadoopS3Keys.
@github-actions github-actions Bot added bug Something isn't working area:writer Native Parquet writer area:Iceberg labels Oct 8, 2026
@manuzhang

manuzhang commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

LGTM overall. Some minor questions.

  1. Do we plan to ignore more common configs?
  2. Do we need to update user and contributor guide as well?

@andygrove

Copy link
Copy Markdown
Member Author

Not a broader list, no. I've only added keys that S3A wouldn't apply to an Iceberg data-file write and that arrive as cluster-wide defaults rather than a per-job choice: Spark 4.1 sets the two committer keys itself whenever spark-hadoop-cloud is on the classpath, and the other three are common platform-level tuning. Anything S3A does apply to a write that the native writer never sees (encryption, ACLs, storage class, signing, the credentials provider) has to keep falling back, which is what #6441 is for. I'd rather add a key when someone hits it blocking native writes than guess at the long tail. Is there one you've run into?

And yes to the guides. Both name the ignored keys (vectored-read and downgrade.syncable.exceptions), so they're stale after this change. I put the update for both in a separate docs-only PR, #6828, to avoid re-running CI here. It should merge after this one.

@andygrove
andygrove added this pull request to the merge queue Oct 9, 2026
@andygrove

Copy link
Copy Markdown
Member Author

Thanks @manuzhang @snmvaughan

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 9, 2026
@andygrove
andygrove added this pull request to the merge queue Oct 10, 2026
Merged via the queue into apache:main with commit a3b297c Oct 10, 2026
41 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:Iceberg area:writer Native Parquet writer bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants