OntoLLMJudge: A Framework for Neurosymbolic Evaluation of LLM Generations
Gendron, Tsaneva & Sabou, 2026
In The 4th Workshop on Evaluation of Language Models in Knowledge Engineering (ELMKE) co-located with ISWC 2026, Bari (Italy), October 2026.
Abstract
Large language models (LLMs) are increasingly utilized to support a wide range of tasks, often open-ended in nature and requiring specialized domain knowledge. However, evaluating outputs, generated by LLMs on such tasks, presents a fundamental trade-off: reliable assessment requires expert judgment, yet manual expert annotation does not scale. Existing automated approaches, such as reference-based metrics and LLM-as-a-judge paradigms, fail to capture actual quality of LLM outputs or inherit the opacity and bias of the models they assess. To address this trade-off, we propose OntoLLMJudge – a neurosymbolic framework that utilizes ontology engineering as a principled control layer over neural evaluation. Rather than leaving expert criteria implicit in prompts or delegating them entirely to a learned model, OntoLLMJudge relies on explicit criteria formalized in an ontology. This makes the evaluation criteria machine-readable, transparent, and independently adaptable by domain experts without modification of the prediction engine. The ontology thus serves a dual role: it informs the neural predictors that score the generated content, and captures the expert-defined criteria for the final score interpretation. We instantiate and validate OntoLLMJudge on the automated evaluation of LLM-generated explanations of ontology modeling mistakes. Our findings indicate that ontology-grounded fine-tuning of scoring predictors consistently improves classification performance over pre-trained models, leading to an average improvement in Macro F1-score across dimensions of 0.49 over pre-trained models and 0.10 over classically fine-tuned models.
Keywords : Knowledge Engineering, LLM Evaluation, LLM-as-a-judge, Neurosymbolic AI
