Apache Arrow
E1312494
UNEXPLORED
Apache Arrow is an open-source, columnar in-memory data format and computing framework designed for high-performance analytics and efficient data interchange across different systems and languages.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Apache Arrow canonical | 3 |
How this entity was disambiguated
This entity first appeared as the object of triple T18205422 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Apache Arrow Context triple: [Hugging Face Datasets, usesFormat, Apache Arrow]
-
A.
Apache Parquet
Apache Parquet is a columnar storage file format optimized for efficient data compression and query performance in big data processing frameworks such as Apache Hadoop and Apache Spark.
-
B.
Apache Avro
Apache Avro is a data serialization system and file format in the Apache Hadoop ecosystem that provides compact, fast, binary data encoding with rich schema support and dynamic typing.
-
C.
Apache Iceberg
Apache Iceberg is an open table format for huge analytic datasets that enables reliable, high-performance querying and data management in data lake environments.
-
D.
Apache Tez
Apache Tez is a distributed data processing framework designed for building high-performance batch and interactive data workflows on Hadoop.
-
E.
Apache Impala
Apache Impala is a massively parallel, SQL-on-Hadoop query engine designed for low-latency, interactive analysis of large-scale data stored in distributed systems.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Apache Arrow Target entity description: Apache Arrow is an open-source, columnar in-memory data format and computing framework designed for high-performance analytics and efficient data interchange across different systems and languages.
-
A.
Apache Parquet
Apache Parquet is a columnar storage file format optimized for efficient data compression and query performance in big data processing frameworks such as Apache Hadoop and Apache Spark.
-
B.
Apache Avro
Apache Avro is a data serialization system and file format in the Apache Hadoop ecosystem that provides compact, fast, binary data encoding with rich schema support and dynamic typing.
-
C.
Apache Iceberg
Apache Iceberg is an open table format for huge analytic datasets that enables reliable, high-performance querying and data management in data lake environments.
-
D.
Apache Tez
Apache Tez is a distributed data processing framework designed for building high-performance batch and interactive data workflows on Hadoop.
-
E.
Apache Impala
Apache Impala is a massively parallel, SQL-on-Hadoop query engine designed for low-latency, interactive analysis of large-scale data stored in distributed systems.
- F. None of above. chosen
Referenced by (3)
Full triples — surface form annotated when it differs from this entity's canonical label.