TY - GEN
T1 - TREX
T2 - 19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026
AU - Won, Inho
AU - Yoo, Hangyeol
AU - Cho, Minkyung
AU - Park, Jungyeul
AU - Song, Hoyun
AU - Lim, Kyung Tae
N1 - Publisher Copyright:
© 2026 Association for Computational Linguistics.
PY - 2026
Y1 - 2026
N2 - Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer’s compression performance critically affects the efficiency of LLM training and inference, existing approaches rely on heuristics or costly large-scale searches to determine optimal language ratios. We introduce Tokenizer Regression for Optimal Data MiXture (TREX), a regression-based framework that efficiently predicts the optimal data mixture for tokenizer training. TREX trains small-scale proxy tokenizers on random mixtures, gathers their compression statistics, and learns to predict compression performance from data mixtures. This learned model enables scalable mixture search before large-scale tokenizer training, mitigating the accuracy-cost trade-off in multilingual tokenizer design. Tokenizers trained with TReX’s predicted mixtures outperform mixtures based on LLaMA3 and uniform distributions by up to 12% in both in- and out-of-distribution compression efficiency, demonstrating strong scalability, robustness, and practical effectiveness.
AB - Building effective tokenizers for multilingual Large Language Models (LLMs) requires careful control over language-specific data mixtures. While a tokenizer’s compression performance critically affects the efficiency of LLM training and inference, existing approaches rely on heuristics or costly large-scale searches to determine optimal language ratios. We introduce Tokenizer Regression for Optimal Data MiXture (TREX), a regression-based framework that efficiently predicts the optimal data mixture for tokenizer training. TREX trains small-scale proxy tokenizers on random mixtures, gathers their compression statistics, and learns to predict compression performance from data mixtures. This learned model enables scalable mixture search before large-scale tokenizer training, mitigating the accuracy-cost trade-off in multilingual tokenizer design. Tokenizers trained with TReX’s predicted mixtures outperform mixtures based on LLaMA3 and uniform distributions by up to 12% in both in- and out-of-distribution compression efficiency, demonstrating strong scalability, robustness, and practical effectiveness.
UR - https://www.scopus.com/pages/publications/105040526335
U2 - 10.18653/v1/2026.eacl-long.298
DO - 10.18653/v1/2026.eacl-long.298
M3 - Conference contribution
AN - SCOPUS:105040526335
T3 - EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers)
SP - 6353
EP - 6370
BT - Long Papers
A2 - Demberg, Vera
A2 - Inui, Kentaro
A2 - Marquez Villodre, Lluis
PB - Association for Computational Linguistics (ACL)
Y2 - 24 March 2026 through 29 March 2026
ER -