Aiyappa et al., "Can we trust the evaluation on ChatGPT?" (2023)
2023-03-22 → 2026-08-14
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn, Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing, 47–54 (2023)
Link | arXiv | PDF
@inproceedings{aiyappa2023can,
author = {Rachith Aiyappa and Jisun An and Haewoon Kwak and Yong-Yeol Ahn},
title = {Can we trust the evaluation on ChatGPT?},
booktitle = {{Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023)}},
pages = {47--54},
address = {Toronto, Canada},
publisher = {Association for Computational Linguistics},
month = {7},
archivePrefix = {arXiv},
eprint = {2303.12767},
primaryClass = {cs.CL},
url = {https://aclanthology.org/2023.trustnlp-1.5/},
doi = {10.18653/v1/2023.trustnlp-1.5},
year = {2023},
}
Examines why closed, continuously updated language models are difficult to evaluate fairly. A stance-detection case study highlights how training-data contamination can compromise ChatGPT evaluations.