Aiyappa et al., "Can we trust the evaluation on ChatGPT?" (2023)

2023-03-22 → 2026-08-14

Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn, Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing, 47–54 (2023)
Link | arXiv | PDF

@inproceedings{aiyappa2023can,
    author = {Rachith Aiyappa and Jisun An and Haewoon Kwak and Yong-Yeol Ahn},
    title = {Can we trust the evaluation on ChatGPT?},
    booktitle = {{Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023)}},
    pages = {47--54},
    address = {Toronto, Canada},
    publisher = {Association for Computational Linguistics},
    month = {7},
    archivePrefix = {arXiv},
    eprint = {2303.12767},
    primaryClass = {cs.CL},
    url = {https://aclanthology.org/2023.trustnlp-1.5/},
    doi = {10.18653/v1/2023.trustnlp-1.5},
    year = {2023},
}

Examines why closed, continuously updated language models are difficult to evaluate fairly. A stance-detection case study highlights how training-data contamination can compromise ChatGPT evaluations.

Data leakage

Receive my updates

YY's Random Walks — Science, academia, and occasional rabbit holes.

YY's Bike Shed — Sustainable mobility, urbanism, and the details that matter.

×