Step-Level Self-Consistency Group Relative Policy Optimization for LLM Reasoning Hallucinations
Researchers introduce SSC-GRPO, a method that assigns step-level rewards to reasoning traces in large language models by computing self-consistency scores across multiple rollouts. This approach targets context-sensitive factual hallucinations, where models possess the necessary knowledge but make errors due to contextual interference. SSC-GRPO demonstrates state-of-the-art performance on mathematical reasoning benchmarks and hallucination leaderboards.
Why it matters: Improving the detection and mitigation of hallucinations in LLM reasoning is crucial for deploying these models in complex, multi-step tasks.
Full story at: arXiv Computation and Language ↗