MAEB
Audio embedding quality across both audio-only and audio-text cross-modal tasks, spanning retrieval, classification, clustering, multilabel classification, pair classification, reranking, and zero-shot classification. Currently in beta pending peer review.
Cite this benchmark
Citation (BibTeX)
@misc{assadi2026maebmassiveaudioembedding,
archiveprefix = {arXiv},
author = {Adnan El Assadi and Isaac Chung and Chenghao Xiao and Roman Solomatin and Animesh Jha and Rahul Chand and Silky Singh and Kaitlyn Wang and Ali Sartaz Khan and Marc Moussa Nasser and Sufen Fong and Pengfei He and Alan Xiao and Ayush Sunil Munot and Aditya Shrivastava and Artem Gazizov and Niklas Muennighoff and Kenneth Enevoldsen},
eprint = {2602.16008},
primaryclass = {cs.SD},
title = {MAEB: Massive Audio Embedding Benchmark},
url = {https://arxiv.org/abs/2602.16008},
year = {2026},
}Languages 159
Tasks 30
Task Types 7
Models 0