Apache Iceberg
E1272509
UNEXPLORED
Apache Iceberg is an open table format for huge analytic datasets that enables reliable, high-performance querying and data management in data lake environments.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Apache Iceberg canonical | 2 |
How this entity was disambiguated
This entity first appeared as the object of triple T17498847 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Apache Iceberg Context triple: [Amazon Athena, supportsDataFormat, Apache Iceberg]
-
A.
Apache Parquet
Apache Parquet is a columnar storage file format optimized for efficient data compression and query performance in big data processing frameworks such as Apache Hadoop and Apache Spark.
-
B.
Delta Lake storage layer
Delta Lake storage layer is an open-source data storage framework that brings ACID transactions, schema enforcement, and reliability to data lakes, particularly in big data and analytics environments.
-
C.
Apache Gobblin
Apache Gobblin is an open-source distributed data integration framework designed for large-scale data ingestion, replication, and lifecycle management across diverse data sources and sinks.
-
D.
Apache Hive
Apache Hive is a data warehouse and SQL-like query system built on top of Hadoop for managing and analyzing large datasets stored in distributed storage.
-
E.
Apache Tez
Apache Tez is a distributed data processing framework designed for building high-performance batch and interactive data workflows on Hadoop.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Apache Iceberg Target entity description: Apache Iceberg is an open table format for huge analytic datasets that enables reliable, high-performance querying and data management in data lake environments.
-
A.
Apache Parquet
Apache Parquet is a columnar storage file format optimized for efficient data compression and query performance in big data processing frameworks such as Apache Hadoop and Apache Spark.
-
B.
Delta Lake storage layer
Delta Lake storage layer is an open-source data storage framework that brings ACID transactions, schema enforcement, and reliability to data lakes, particularly in big data and analytics environments.
-
C.
Apache Gobblin
Apache Gobblin is an open-source distributed data integration framework designed for large-scale data ingestion, replication, and lifecycle management across diverse data sources and sinks.
-
D.
Apache Hive
Apache Hive is a data warehouse and SQL-like query system built on top of Hadoop for managing and analyzing large datasets stored in distributed storage.
-
E.
Apache Tez
Apache Tez is a distributed data processing framework designed for building high-performance batch and interactive data workflows on Hadoop.
- F. None of above. chosen
Referenced by (2)
Full triples — surface form annotated when it differs from this entity's canonical label.