A survey on LLM-as-a-Judge
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains challenging due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success, leading to ”LLM-as-a-Judge,” where LLMs serve as evaluators for complex tasks. With their ability to process diverse data types and provide scalable assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge requiring careful design and standardization.
This paper provides a comprehensive survey on LLM-as-a-Judge, offering a formal definition and detailed classification, while addressing the core question: How to build reliable LLM-as-a-Judge systems? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse scenarios. We propose methodologies for evaluating reliability, supported by a novel benchmark. To advance development and deployment, we discuss practical applications, challenges, and future directions. Our contributions span multiple levels: we establish conceptual boundaries, reorganize fragmented literature into a unified framework, and propose a reliability-oriented benchmark. We articulate a forward-looking research agenda, offering theoretical foundations and practical guidance for constructing reliable and trustworthy LLM-as-a-Judge systems. Resources are available at https://awesome-llm-as-a-judge.github.io/. Large Language Models, LLM-as-a-Judge, Automated Evaluation, Reliability Assessment, Trustworthy AI
