An Item Response Theory-based R Module for Algorithm Portfolio Analysis

doi:10.48550/arXiv.2408.14025

An Item Response Theory-based R Module for Algorithm Portfolio Analysis

Experimental evaluation is crucial in AI research, especially for assessing algorithms across diverse tasks. Many studies often evaluate a limited set of algorithms, failing to fully understand their strengths and weaknesses within a comprehensive portfolio. This paper introduces an Item Response Theory (IRT) based analysis tool for algorithm portfolio evaluation called AIRT-Module. Traditionally used in educational psychometrics, IRT models test question difficulty and student ability using responses to test questions. Adapting IRT to algorithm evaluation, the AIRT-Module contains a Shiny web application and the R package airt. AIRT-Module uses algorithm performance measures to compute anomalousness, consistency, and difficulty limits for an algorithm and the difficulty of test instances. The strengths and weaknesses of algorithms are visualised using the difficulty spectrum of the test instances. AIRT-Module offers a detailed understanding of algorithm capabilities across varied test instances, thus enhancing comprehensive AI method assessment. It is available at https://sevvandi.shinyapps.io/AIRT/ .

Publication:

arXiv e-prints

Pub Date:

August 2024

DOI:

10.48550/arXiv.2408.14025

arXiv:

arXiv:2408.14025

Bibcode:

2024arXiv240814025O

Keywords:

Computer Science - Machine Learning

E-Print:

10 Pages, 6 Figures. Submitted to SoftwareX

NASA/ADS

An Item Response Theory-based R Module for Algorithm Portfolio Analysis

Abstract