Main Article Content

The assessment of academic theses by human examiners is traditionally considered the gold standard. This article challenges this assumption, discusses typical strengths and weaknesses of human evaluations, and contrasts them with assessments of six AI models. Based on 52 bachelor's theses, statistical measures of similarity as well as the qualitative depth of the reports are compared. The results show that human examiners utilize the grading scale more broadly and tend to grade more leniently, while AI models lean toward leveling. In qualitative performance assessment, AI models systematically identify complementary aspects—a strong argument for hybrid human-AI approaches.

Article Details