A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios

doi:10.48550/arXiv.2408.01963

A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios

We evaluate the robustness of several large language models on multiple datasets. Robustness here refers to the relative insensitivity of the model's answers to meaning-preserving variants of their input. Benchmark datasets are constructed by introducing naturally-occurring, non-malicious perturbations, or by generating semantically equivalent paraphrases of input questions or statements. We further propose a novel metric for assessing a model robustness, and demonstrate its benefits in the non-adversarial scenario by empirical evaluation of several models on the created datasets.

Publication:

arXiv e-prints

Pub Date:

August 2024

DOI:

10.48550/arXiv.2408.01963

arXiv:

arXiv:2408.01963

Bibcode:

2024arXiv240801963A

Keywords:

Computer Science - Computation and Language;
Statistics - Applications

E-Print:

Published in the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP) findings

NASA/ADS

A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios

Abstract