← 返回论文检索
ICML 2026PosterAccept (regular)

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

Peihao Wang, Shan Yang, Xijun Wang, Tesi Xiao, Jiahui Gao, Changlong Yu, Yu Lou, Pan Li, Zhangyang “Atlas” Wang, Ming Lin, Rene Vidal

The University of Texas at Austin · Adobe Foundry · Amazon · Hong Kong University of Science and Technology · Department of Computer Science and Engineering, The Hong Kong University of Science and Technology · Georgia Institute of Technology · XTX Markets & University of Texas at Austin · UMD-CP & UNC-CH · University of Pennsylvania & Amazon

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by *projecting future states and selecting goal-directed actions*, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learning or test-time training, planning remains external to the model architecture. We formulate reasoning as *optimal control* and introduce the *Test-Time Control (TTC)* layer, which performs finite-horizon *LQR* planning over latent states at inference time, enabling *planning before prediction*. To ensure scalability, we derive a hardware-efficient LQR solver based on a symplectic formulation and implement it as a fused CUDA kernel, enabling parallel execution with minimal overhead. Integrated as an adapter into pretrained LLMs, TTC layers improve mathematical reasoning performance by up to +27.8 on MATH-500 and 2-3x Pass@8 improvements on AMC and AIME, demonstrating that embedding optimal control as an architectural component provides an effective and scalable mechanism for reasoning beyond test-time training.