TY - GEN
T1 - Beyond Accuracy
T2 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
AU - Park, Jieun
AU - Lim, Kyung Tae
AU - Lim, Joon Ho
N1 - Publisher Copyright:
©2026 Association for Computational Linguistics.
PY - 2026
Y1 - 2026
N2 - Recent advancements in LLMs have significantly improved mathematical problem-solving, with models like GPT-4 achieving human-level performance. However, proficiently solving mathematical problems differs fundamentally from effectively teaching mathematics. To bridge this gap, we introduce the Bi-GSM8K benchmark, a bilingual English-Korean dataset enriched with teacher solutions, student solutions, and annotations marking students’ initial errors. This dataset is designed to evaluate two core capabilities of LLMs: (1) measuring similarity between student and teacher solutions, and (2) identifying the initial error point in student solutions. Our method achieves high agreement with human judgments, with Pearson 0.89 and Spearman 0.88 on English, and Pearson 0.89 and Spearman 0.87 on Korean. It also offers significantly lower latency and resource usage than commercial APIs, demonstrating strong computational efficiency. In the error detection task, open-source models achieved approximately 86% accuracy, with performance within 10% points of commercial LLMs API, suggesting strong practical potential. Our key contributions include the open-source release of Bi-GSM8K, novel evaluation metrics, and comparative analyses of LLM performance across languages.
AB - Recent advancements in LLMs have significantly improved mathematical problem-solving, with models like GPT-4 achieving human-level performance. However, proficiently solving mathematical problems differs fundamentally from effectively teaching mathematics. To bridge this gap, we introduce the Bi-GSM8K benchmark, a bilingual English-Korean dataset enriched with teacher solutions, student solutions, and annotations marking students’ initial errors. This dataset is designed to evaluate two core capabilities of LLMs: (1) measuring similarity between student and teacher solutions, and (2) identifying the initial error point in student solutions. Our method achieves high agreement with human judgments, with Pearson 0.89 and Spearman 0.88 on English, and Pearson 0.89 and Spearman 0.87 on Korean. It also offers significantly lower latency and resource usage than commercial APIs, demonstrating strong computational efficiency. In the error detection task, open-source models achieved approximately 86% accuracy, with performance within 10% points of commercial LLMs API, suggesting strong practical potential. Our key contributions include the open-source release of Bi-GSM8K, novel evaluation metrics, and comparative analyses of LLM performance across languages.
UR - https://www.scopus.com/pages/publications/105039193030
U2 - 10.18653/v1/2026.findings-eacl.85
DO - 10.18653/v1/2026.findings-eacl.85
M3 - Conference contribution
AN - SCOPUS:105039193030
T3 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
SP - 1678
EP - 1704
BT - 19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
PB - Association for Computational Linguistics (ACL)
Y2 - 24 March 2026 through 29 March 2026
ER -