Biologically grounded locality priors close the data gap for vision transformers in neural prediction

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

For datasets with thousands of neurons and images, vision transformers have proven successful at predicting neural responses to stimuli. However, they are expected to underperform in low-data regimes, where CNNs and Gaussian processes are considered more effective. We ask whether transformers can be made competitive for small-scale neural prediction, and show that underperformance in this regime can be overturned with the right inductive bias. We equip a vision transformer with a differentiable per-neuron circular crop in feature space. The crop is centered on each neuron’s receptive field, with a radius selected per neuron, so the model only sees the small image region that drives that neuron instead of the whole image. This makes the cost of attention scale with the size of the receptive field rather than with the size of the image. A single transformer stack is shared across all neurons: each neuron’s specificity resides in the crop, not in the architecture. We call this model circular receptive-field vision transformer (CiRF-ViT). We evaluate it on small multi-electrode-array recordings of mouse and salamander retinas: a mouse preparation of 41 ganglion cells and two salamander preparations totaling 49 ganglion cells, each with only a few thousand stimulus–response pairs, two orders of magnitude below the scale at which transformers are typically trained. Against CNN and Gaussian-process baselines, CiRF-ViT reaches the highest mean explained variance on both datasets (0.94 on mouse, 0.95 on salamander). Probed with the local spike-triggered average (LSTA), a zero-shot test of context-dependent, nonlinear stimulus sensitivity, CiRF-ViT reproduces the qualitative polarity inversion that a linear model cannot capture by construction. A biologically grounded, per-neuron locality prior is therefore enough to make transformers competitive for neural prediction well below their usual data scale, while matching the reference models on an established functional signature of the retina’s nonlinear computation.

Article activity feed