All registries
Meta-indexIn Rundown bodies

Apache Software Foundation Projects

Apache ASF

Cloud & APIsData & file formatsMeta-indexes
WHY YOU NEED THIS

Data infrastructure specs (Parquet, Avro, Arrow) live under Apache — broader than the Parquet/Arrow entry alone.

What It Is

Catalog of 300+ Apache projects: Kafka, Spark, Hadoop, Cassandra, HTTP Server, Tomcat, and formal specs for Parquet, Arrow, Avro, Thrift.

Taxonomy

By category: Big Data · Web · Libraries · Incubator · Attic

Top-Level Categories

Big Data (Kafka, Spark, Hadoop)Web servers (httpd, Tomcat)Formats (Parquet, Avro, Arrow)Lucene/Solr

Master Catalog

Sub-Indexes & Guides

Rundown Standards Body
Apache Software Foundation

Open-source foundation publishing columnar and analytics file formats widely used in data pipelines: Apache Parquet (columnar storage) and Apache Arrow (in-memory columnar + IPC serialization). Specs are maintained on parquet.apache.org and arrow.apache.org.

View all bodies →
Visit

Curated specs from Apache ASF

Apache ParquetApacheShould Know

Parquet

Parquet is the standard lakehouse/warehouse on-disk format (Spark, DuckDB, BigQuery export, Snowflake). Column chunk layout explains why analytics workloads beat row-oriented JSON/CSV.

ProductBack OfficeData Formats
Details
Apache ArrowApacheNiche

Arrow IPC

Arrow is the interchange layer between Python/R/JVM analytics, DuckDB, and Parquet readers. IPC framing is how columnar data crosses process boundaries without JSON overhead.

ProductBack OfficeData Formats
Details

Related Registry Sources

apachekafkasparkavroopen-source