JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models
JOR-Bench is a suite of five Japanese-language benchmarks designed to evaluate large language models (LLMs) on operations research (OR) problems, comprising 1,319 problems across linear, mixed-integer, non-linear, and combinatorial optimization. The benchmarks are Japanese translations of established English datasets, enabling direct cross-lingual comparison. Evaluation of seven LLMs shows that strong multilingual models exhibit nearly identical OR formulation accuracy in Japanese and English, with only a -0.3 percentage point difference, though error analysis uncovers subtle language-specific failure modes.
Why it matters: JOR-Bench provides a rigorous tool for assessing LLMs' mathematical reasoning in Japanese, highlighting both the strengths and nuanced limitations of multilingual models across languages.
Full story at: arXiv Computation and Language ↗