MTEB logo MTEB
dense Open weights ST compatible text

BAAI/bge-m3

Languages
AfrikaansAmharicAsturianBelarusianBengaliBulgarianCatalanCebuanoCentral KurdishChineseDanishEnglish +16 more Show fewer EstonianFinnishFrenchGalicianGermanGujaratiHebrewHindiItalianJapaneseKoreanModern Greek (1453-)North AzerbaijaniRussianThaiUkrainian
License
mit
Trained on
CMedQAv1-rerankingCMedQAv2-rerankingCodeSearchNetDuRetrievalHotpotQAHotpotQA-NLHotpotQA-PLHotpotQAHardNegatives +18 more Show fewer LeCaRDv2MIRACLRerankingMIRACLRetrievalMIRACLRetrievalHardNegativesMMarcoRerankingMSMARCOMSMARCO-PLMSMARCOHardNegativesMrTidyRetrievalNQNQ-NLNQ-PLNQHardNegativesNanoMSMARCORetrievalNanoNQRetrievalT2RerankingT2RetrievalmMARCO-NL
Cite this model
Citation (BibTeX)
@misc{bge-m3,
      title={BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
      author={Jianlv Chen and Shitao Xiao and Peitian Zhang and Kun Luo and Defu Lian and Zheng Liu},
      year={2024},
      eprint={2402.03216},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
Parameters 568M
Active parameters 312M
Embedding dim 1,024
Max tokens 8,192
Memory 2.1 GB
Released 2024-06-28

Openness

  • Open weights
  • Open license
  • Training code
  • Training data
  • Paper
  • Model card

Benchmark scores

Loading…