A theory of multi-task computation and task selection
Curation statements for this article:-
Curated by eLife
eLife Assessment
This valuable study analyzes how recurrent neural network can flexibly switch between multiple tasks. The evidence is solid, but reviewers have raised questions about the impacts of transient dynamics, whether the activity is actually chaotic, can inputs be used to switch between tasks, what determines overlap between tasks and others. In addition, there were more minor questions pertaining to whether the analyses are actually analytic or primarily based on numerical simulations, as well as the relationship to prior studies where the same network was optimized to perform multiple tasks simultaneously with different readouts.
This article has been Reviewed by the following groups
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
- Evaluated articles (eLife)
Abstract
Neural activity during the performance of a stereotyped behavioral task is often described as low-dimensional, occupying only a limited region in the space of all firing-rate patterns. This region has been referred to as the “neural manifold” associated with a task. More recently, recordings of neural activity in animals challenged to perform multiple tasks have suggested that each task is associated with a different low-dimensional manifold. What connectivity structures underlie this flexibility in neural dynamics, and how is interference between the dynamics associated with different tasks avoided? We develop a theoretical model for multi-task computation in nonlinear recurrent neural networks whose connectivity is constructed as a weighted sum of many low-rank components, each encoding the dynamics associated with a different task. The model demonstrates that interference between different tasks’ dynamics limits flexible multi-tasking and can lead to chaotic fluctuations. However, small modulations of a network’s effective connectivity overcome this interference. We derive the conditions that enable such task selection and characterize both single-neuron and population statistics in task-selected and unselected states. The model reveals the requirements for a single network to produce distinct dynamics confined to distinct neural manifolds and suggests circuit mechanisms that support this capability. Using the model, we propose different hypotheses for explaining the origin of high-dimensional neural activity in large-scale recordings.
Article activity feed
-
eLife Assessment
This valuable study analyzes how recurrent neural network can flexibly switch between multiple tasks. The evidence is solid, but reviewers have raised questions about the impacts of transient dynamics, whether the activity is actually chaotic, can inputs be used to switch between tasks, what determines overlap between tasks and others. In addition, there were more minor questions pertaining to whether the analyses are actually analytic or primarily based on numerical simulations, as well as the relationship to prior studies where the same network was optimized to perform multiple tasks simultaneously with different readouts.
-
Reviewer #1 (Public review):
Summary:
Marschall et al. develop a theoretical framework for analyzing multi-task dynamics in nonlinear recurrent neural networks (RNNs). In the RNN model, recurrent connectivity is a linear superposition of multiple non-overlapping low-rank components, each corresponding to a separate "task". Each task is an autonomous dynamical system that does not incorporate external inputs. The "multi-task computation" setting examines whether multiple tasks (dynamical systems) can operate concurrently or how a network can switch between tasks.
Within this framework, the authors show that when connectivity consists of two low-rank components implementing two tasks (a limit cycle and a bistable attractor), the network exhibits winner-takes-all dynamics, resulting in only one task being active and the other one …
Reviewer #1 (Public review):
Summary:
Marschall et al. develop a theoretical framework for analyzing multi-task dynamics in nonlinear recurrent neural networks (RNNs). In the RNN model, recurrent connectivity is a linear superposition of multiple non-overlapping low-rank components, each corresponding to a separate "task". Each task is an autonomous dynamical system that does not incorporate external inputs. The "multi-task computation" setting examines whether multiple tasks (dynamical systems) can operate concurrently or how a network can switch between tasks.
Within this framework, the authors show that when connectivity consists of two low-rank components implementing two tasks (a limit cycle and a bistable attractor), the network exhibits winner-takes-all dynamics, resulting in only one task being active and the other one suppressed. A task consistently dominates this competition when the magnitude of its low-rank connectivity ("task strength") exceeds that of the other task. When one task is dominant and many other tasks with weaker low-rank connectivity are present, increasing the number of tasks destabilizes the dominant-task dynamics, leading to chaotic fluctuations in network activity. Supported by the dynamical mean-field theory analysis, the authors show that as the dominant-task strength increases, the network transitions through three dynamical regimes: chaotic spontaneous activity, chaotic task-selected dynamics, and non-chaotic task-selected dynamics. The theoretical analysis additionally predicts how latent task-related dynamics manifest in single-neuron activity and how the dimensionality of population activity changes across the three dynamical regimes.
The results are interesting, the analyses and simulations are rigorous, and the text is clear and easy to follow. Overall, this study is a significant and timely contribution to the literature on low-rank RNNs, an influential model class for low-dimensional neural dynamics in computational neuroscience.
Main comments:
(1) In this modeling framework, only one dominant task can be selected while all other tasks are suppressed. In contrast, several previous studies constructed RNNs (either through gradient-descent optimization or reservoir computing) that simultaneously generate outputs for multiple tasks across the corresponding task-specific readouts. Of course, what counts as a task is arbitrary, and one could view the dynamics of a reservoir network as implementing a single high-dimensional "task" with multiple readouts. Nevertheless, it would be helpful to explicitly clarify the distinction and similarities between the current modeling framework and networks that simultaneously solve multiple tasks.
(2) The role of external inputs in task selection appears to be underdeveloped. It is only briefly examined in Fig. S4, with the conclusion that external inputs aligned with the task subspace cannot enable selection of the desired task. However, previous multi-task RNN models (e.g., optimized through gradient descent) are clearly able to switch across many tasks using external inputs. In these models, external inputs modulate RNN activity along specific directions, shaped through gradient descent, to select the relevant task for each input. In contrast, this study only considers inputs aligned with the m-direction (left loading vector in the low-rank connectivity for a task). Such input cannot enhance the corresponding task's activity through recurrent amplification (Fig. S4). Yet, low-rank RNN theory predicts that inputs aligned with the n-direction (task's right loading vector) are selectively amplified by the recurrent dynamics. Why inputs aligned with the n-direction were not examined? More broadly, focusing only on inputs aligned with either m- or n-direction appears too narrow, as trained multi-task RNNs indicate that task selection through external inputs is possible, but may require input directions different from m and n.
(3) It is unclear what the reason is for the task interference: overlaps between task loading vectors for different tasks or other nonlinear effects? This issue is especially prominent in the analysis of task capacity (Fig. 2D and Fig. 4A). The text states that the loading vectors are independent across tasks, i.e. there is zero expected overlap between loading vectors of different tasks (line 126). For a network of N neurons, a rank-R task requires 2R loading vectors. Under the assumption of independence, the maximum number of tasks is P = N/(2R). Then α=P/N can be at most 1/(2R). In the simulations, it is chosen R=2, such that alpha can be at most 1/4 for the loading vectors to remain independent. However, the range of alpha reaches up to 1 in Fig. 2D and Fig. 4A, suggesting that loading vectors are no longer independent across tasks for larger alpha. Is the linear dependence between loading vectors (i.e. overlap across tasks) the reason for the task-1 component norm to drop sharply around alpha=1/4 in Fig. 2D? More broadly, is it possible to isolate the contribution of overlap in task loading vectors versus other nonlinear effects?
(4) In the current version of the paper, it is difficult to understand how the analysis based on the dynamical mean-field theory in Methods explains the key observations from the numerical simulations. For example, the mechanism underlying the transition from the spontaneous state to chaotic task-selected state, and to the non-chaotic task-selected state remain opaque. Further, it would be helpful to specify which section in Methods is being referred to at each mention throughout the main text. Finally, Fig. 7 and Eqs. (11-13) clearly state the input sources that drive single unit activity, non-dominant task latent states and dominant task latent state. The decomposition is potentially very informative, but its implications are not discussed in sufficient detail. It would be helpful to provide an intuitive explanation of how contributions of these input sources evolve as connectivity strength of one task increases, and which of them eventually leads to the loss of stability of the previous network state.
(5) The text states that the results can be easily extended to include task-specific inputs and outputs (lines 117-119). However, such an extension does not seem to be straightforward and requires additional explanations. The dynamical mean-field theory analyses here are stationary and describe the steady-state of network dynamics. In contrast, common input-output tasks typically involve transient dynamics, in which time-dependent external inputs keep changing the RNN flow field, and steady-state is never reached. Under these transient conditions, it is unclear whether the same conclusions apply. For example, if a network is at a low-activity baseline when a task-input begins to drive activity in the corresponding task subspace, it is unclear whether the activity in other task-subspaces would grow sufficiently fast to cause interference, or whether such interference would not be observed. Thus, a more detail analyses are necessary to support the extension of the results to input-driven transient tasks beyond autonomous dynamical systems.
(6) When task strength is the same for all tasks, what determines which task will win the competition? Is it frozen noise in connectivity such that one task always wins, or do initial conditions determine which task wins, based on which task's activity grows faster?
(7) Does the theory require the activity of all neurons to operate in the saturating part of nonlinearity? For example, the text states "increased activity reduces the gain factor <Φ'(t)>" (line 169). This statement is only true when Φ'<0. For sigmoid nonlinearity used in the paper, Φ'>0 when firing rate is small. If a substantial fraction of neurons in the network is near the rest state, would the theory still apply? Similarly, this statement does not hold for ReLU nonlinearity, and it is unclear how the theory applies to ReLU networks in Fig. S3. The paper states that the results are not specific to the choice of nonlinearity (lines 176-178). However, the dynamics being studied (bistability and limit cycles) both operate on the saturating part of the nonlinearity. Could the authors clarify the assumptions on the nonlinearity for the theory to apply?
(8) Is it possible to interpret the results in Fig. 4? What does this dependence on the overlap matrix mean? Is there an intuition for this particular dependence, or is it just an observation without general interpretation?
(9) The results in Fig. 5E appear underdeveloped and somewhat arbitrary. It is unclear how the dependence of dimensionality on the recording time would change as a function of time spent in a task. If this time is long, then the curve grows slowly and total dimensionality is high. If this time is very short, then the network may not have sufficient time for all task variables to grow sufficiently large to contribute significantly to the total variance. Thus, the grows may be faster and the total variance may saturate at a lower value. Hence, it is unclear whether there will be always a qualitative difference from the spontaneous activity curve. Furthermore, since only one of two curves is measured, what quantitative criteria should be used to determine whether it is consistent with task switching or spontaneous state?
(10) On line 368: "For sufficiently large number of tasks, the dimensionality associated with sequential task selection can greatly exceed that of the spontaneous state (Fig. 5E inset)" - it seems that Fig. 5E inset shows the opposite that the dimensionality of spontaneous state can saturate at a very high value for large N, exceeding the dimension of task-switching network in the main plot. Although it is hard to say, since the inset has many lines with only two labels and no ticks on axis, so it is unclear what exactly does it show.
(11) Related to discussion on line 368-370: In a task-switching state, would the switching between tasks also be reflected in behavior? In addition, the time-correlation functions would not be stationary in task-switching state, i.e. they would change over time, whereas they will be stationary in the spontaneous state. Thus, could the two mechanisms be dissociated in experiments using behavior or metrics beyond dimensionality?
(12) On line 385, could the authors provide more details on how they envision the two potential mechanisms-synaptic plasticity and targeted neuromodulation-to reinforce a task-specific low-rank connectivity pattern? If neuromodulation changes the gain of individual neurons, this modulation corresponds to scaling of connectivity by a diagonal matrix, not strengthening of a specific low-rank component. Short-term facilitation or depression also modulates synaptic strength depending on the activity of the presynaptic neuron, thus also scaling connectivity by a diagonal matrix. It is unclear how these two mechanisms could produce a specific low-rank modulation.
-
Reviewer #2 (Public review):
Summary
The authors ask what recurrent connectivity supports many distinct task-related manifolds when the associated dynamics interfere, how a circuit engages one task while suppressing others, and what produces high-dimensional activity. Extending previous theoretical studies on low-dimensional dynamics in large networks, they use a solvable model whose weight matrix is a weighted sum of many low-rank, task-specific components and develop a dynamical mean-field theory that relates connectivity, dynamics, and measurable population signatures of multi-tasking.
Strengths
(1) The question is timely. Low-rank networks are a leading model for low-dimensional latent dynamics, and the composition of dynamical systems has been proposed as a mechanism allowing for rapid, flexible learning; the paper connects these …
Reviewer #2 (Public review):
Summary
The authors ask what recurrent connectivity supports many distinct task-related manifolds when the associated dynamics interfere, how a circuit engages one task while suppressing others, and what produces high-dimensional activity. Extending previous theoretical studies on low-dimensional dynamics in large networks, they use a solvable model whose weight matrix is a weighted sum of many low-rank, task-specific components and develop a dynamical mean-field theory that relates connectivity, dynamics, and measurable population signatures of multi-tasking.
Strengths
(1) The question is timely. Low-rank networks are a leading model for low-dimensional latent dynamics, and the composition of dynamical systems has been proposed as a mechanism allowing for rapid, flexible learning; the paper connects these two ideas under a single theoretical framework.
(2) The proposal that sequential transitions between low-dimensional, low-rank dynamics can account for the _apparent growth of dimensionality with recording time_ is novel and is the paper's most valuable conceptual contribution.
(3) The mathematical analysis is rigorous, and the spontaneous-state theory is convincingly validated against simulation.
(4) The model produces concrete, falsifiable predictions - heterogeneous, syllable-dependent single-neuron tuning, low within-state dimensionality despite single-neuron variability, and distinct dimensionality-versus-recording-time signatures for the spontaneous versus task-switching accounts.
Weaknesses - whether the claims are supported by the data
(1) Chaos is named but not demonstrated._ The large-P and intermediate task-selected regimes are labeled "chaotic," but the manuscript does not establish chaos. In a homogeneous network, it is known that once the fixed point loses stability, the surviving solution is chaotic (Sompolinsky, Crisanti & Sommers 1988); that guarantee does not transfer here. The DMFT noise term is not computed analytically, and the single-neuron correlation functions (Fig. 5) show disorder, not a demonstrated decay of the fluctuation autocorrelation to zero, nor a positive largest Lyapunov exponent. The concern is sharpened by the possibility of _transient_ chaos: orthogonal to a dominant limit cycle, fluctuations may be locally unstable only at certain amplitudes or phases, so the global attractor could remain a stable cycle visited with chaotic excursions. As it stands, the claim of chaos in the intermediate regime is unsupported; it may well hold for some range of the selected-task strength, but this is neither shown numerically nor proven.
(2) "Analytical theory" overstates what is solved in closed form._ For the task-selected state, the kernels are non-stationary: The DMFT is entrained to the dominant task's dynamics, with an O(1) time-dependent quantity inside the nonlinearity. To my knowledge, there is no closed-form DMFT solution under these conditions. The Methods section supports this, explaining that the general scheme is solved by iterative numerical self-consistency (and described there as prohibitively expensive), and tractability is recovered only in a special block-Haar ensemble with Gaussian currents. This is entirely reasonable, but the main text presents it as an analytical theory; the reliance on numerical solutions of the self-consistency equations should be stated plainly.
(3) The spontaneous-state transition is the classical critical-gain transition, only reparametrized._ The onset of the no-task-dominant state is governed by $g_{eff}^2 = \alpha R\langle D^2\rangle$. It appears to depend on the number of tasks only because per-task strength D is held fixed as tasks accumulate; under a normalization that holds g_eff fixed, the transition reduces to a critical-gain point independent of P, as in extensive random networks. Relatedly, the result that chaos "arises solely from learning many tasks" is, mechanistically, random-network chaos: the random task components raise the weight variance and play the role of effective disorder. This is a legitimate and appealing reframing, but it is not a new transition, and the manuscript should make the relationship to the standard criterion explicit.
(4) Significance of the selection mechanism._ That boosting a task's gain selects it is intuitive, and the authors note the extreme (one $D^\mu$ dominating) is trivial. The non-trivial and genuinely useful contribution is quantitative - that only a small, O(1/P) modulation near criticality is required. This deserves to be foregrounded rather than left to the Discussion.
-