MTEB logo MTEB
retrieval text

CosQA

The dataset is a collection of natural language queries and their corresponding code snippets. The task is to retrieve the most relevant code snippet for a given query.

Domains: ProgrammingWritten
Languages: Englishpython
Task type
Retrieval
Main metric
ndcg_at_10
Source dataset
CoIR-Retrieval/cosqa
License
mit
Dates
2021-05-07 → 2021-05-07
Annotations
derived
Sample creation
found
Cite this task
Citation (BibTeX)

@misc{huang2021cosqa20000webqueries,
  archiveprefix = {arXiv},
  author = {Junjie Huang and Duyu Tang and Linjun Shou and Ming Gong and Ke Xu and Daxin Jiang and Ming Zhou and Nan Duan},
  eprint = {2105.13239},
  primaryclass = {cs.CL},
  title = {CoSQA: 20,000+ Web Queries for Code Search and Question Answering},
  url = {https://arxiv.org/abs/2105.13239},
  year = {2021},
}
Models scored
Dataset preview CoIR-Retrieval/cosqa Open on HuggingFace →

Model scores

Loading…