Weight initialization shapes task organization in recurrent neural networks
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Flexibly recombining computational modules is essential for biological and artificial neural networks to rapidly adapt to changing environments. This requires modules to be shared across tasks rather than rigidly segregated, yet what determines this organization remains unknown. Previous work suggests that weight initialization shapes whether networks learn task-specific or generic representations, but it is unclear whether this extends to recurrent networks and, more importantly, to network connectivity. Here, we systematically vary the initial weight variance of recurrent neural networks and study them using a framework that allows us to identify the functionally relevant connectivity subspaces for each computational module. We find that networks with low initial weight variance converge to solutions in which different subtasks rely on largely overlapping weight subspaces, whereas high-variance networks implement subtasks in higher-dimensional, more segregated weight subspaces. Our results also provide mechanistic insights with implications for interpreting biological neural circuits and for designing efficient recurrent architectures.