← 返回论文检索
ICML 2026PosterAccept (regular)

Efficient and Minimax-optimal In-context Nonparametric Regression with Transformers

Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William Underwood, Richard Samworth

University of Cambridge · ETHZ - ETH Zurich

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We study in-context learning for nonparametric regression with $\alpha$-Hölder smooth regression functions, for some $\alpha>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $\Theta(\log n)$ parameters and $\Omega\bigl(n^{2\alpha/(2\alpha+d)}\log^3 n\bigr)$ pretraining sequences can achieve the minimax-optimal rate of convergence $O\bigl(n^{-2\alpha/(2\alpha+d)}\bigr)$ in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.