Abstract

Continuous attractors offer a unique class of solutions for storing continuousvalued variables in recurrent system states for indefinitely long time intervals.

Unfortunately, continuous attractors suffer from severe structural instability in general—they are destroyed by most infinitesimal changes of the dynamical law that defines them.

This fragility limits their utility especially in biological systems as their recurrent dynamics are subject to constant perturbations.

连续吸引子为在循环系统状态中无限期存储连续值变量提供了一类独特的解决方案.

不幸的是, 连续吸引子通常存在严重的结构不稳定性——它们会被定义它们的动力学规律的大多数微小变化所破坏.

这种脆弱性限制了它们的实用性, 尤其是在生物系统中, 因为它们的循环动力学受到持续微扰.

We observe that the bifurcations from continuous attractors in theoretical neuroscience models display various structurally stable forms. Although their asymptotic behaviors to maintain memory are categorically distinct, their finite-time behaviors are similar.

We build on the persistent manifold theory to explain the commonalities between bifurcations from and approximations of continuous attractors. Fast-slow decomposition analysis uncovers the existence of a persistent slow manifold that survives the seemingly destructive bifurcation, relating the flow within the manifold to the size of the perturbation. Moreover, this allows the bounding of the memory error of these approximations of continuous attractors.

我们观察到理论神经科学模型中连续吸引子的 分岔 显示出各种结构稳定的形式. 尽管它们维持记忆的渐近行为在类别上是不同的, 但它们的有限时间行为是相似的.

我们基于 持久流形理论 来解释连续吸引子的分岔和近似之间的共性. 快慢分解分析揭示了一个持久慢流形的存在, 它在看似破坏性的分岔中幸存下来, 将流动与微扰大小联系起来. 此外, 这允许对这些连续吸引子近似的记忆误差进行界定.

Finally, we train recurrent neural networks on analog memory tasks to support the appearance of these systems as solutions and their generalization capabilities. Therefore, we conclude that continuous attractors are functionally robust and remain useful as a universal analogy for understanding analog memory.

最后, 我们在模拟记忆任务上训练 RNN, 以支持这些系统作为解决方案的出现及其泛化能力. 因此, 我们得出结论, 连续吸引子在功能上是稳健的, 并且仍然有用, 作为理解模拟记忆的普遍类比.

Introduction

Biological systems exhibit robust behaviors that require neural information processing of analog variables such as intensity, direction, and distance. Virtually all neural models of working memory for continuous-valued information rely on persistent internal representations through recurrent dynamics.

The continuous attractor structure in their recurrent dynamics has been a pivotal theoretical tool due to their ability to maintain activity patterns indefinitely through neural population states 1–4. They are hypothesized to be the neural mechanism for the maintenance of eye positions, heading direction, self-location, target location, sensory evidence, working memory, and decision variables, to name a few 5–7. Observations of persistent neural activity across many brain areas, organisms, and tasks have corroborated the existence of continuous attractors.

生物系统表现出稳健的行为, 这些行为需要对 模拟变量 (如强度、方向和距离)进行神经信息处理. 几乎所有用于连续值信息的工作记忆的神经模型都依赖于通过循环动力学实现的持久内部表示.

由于它们能够通过神经群体状态无限期地维持活动模式, 它们在循环动力学中的连续吸引子结构一直是一个关键的理论工具. 人们假设它们是维持 眼睛位置、航向方向、自我位置、目标位置、感官证据、工作记忆和决策变量 等的神经机制. 在许多大脑区域、生物体和任务中观察到的持续神经活动证实了连续吸引子的存在.

Despite their widespread adoption as models of analog memory, continuous attractors are brittle mathematical objects, casting significant doubts on their ontological value and hence suitability in accurately representing biological functions.

Even the smallest arbitrary change in recurrent dynamics can be problematic, destroying the continuum of fixed points essential for continuous-valued working memory. In neuroscience, this vulnerability is well-known and often referred to as the "fine-tuning problem".

尽管它们被广泛采用作为模拟记忆的模型, 但连续吸引子是 脆弱的数学对象, 这对它们的本体价值以及准确表示生物功能的适用性提出了重大质疑.

即使是循环动力学中的最小任意变化也可能是有问题的, 破坏了连续值工作记忆所必需的不动点连续体. 在神经科学中, 这种脆弱性是众所周知的, 通常被称为 "微调问题".

There are two primary sources of perturbations in the recurrent network dynamics:

(1) the stochastic nature of online learning signals that act via synaptic plasticity, and

(2) spontaneous fluctuations in synaptic weights.

Thus, additional mechanisms are necessary to compensate for the degradation in particular implementations, by bringing the short-term behavior closer to that of a continuous attractor.

循环网络动力学中微扰的两个主要来源是:

(1)通过突触可塑性作用的在线学习信号的随机性, 以及

(2)突触权重的自发涨落.

因此, 需要额外的机制来补偿特定实现中的 退化, 通过使 短期行为 更接近连续吸引子的行为.

However, we lack the theoretical basis to understand how much this matters in practice, i.e. what are the effects of different levels of degradation on memory. This is fundamental to justify relying on the brittle concept of continuous attractors for understanding biological analog working memory.

然而, 我们缺乏理论基础来理解这实际上有多重要, 即不同水平退化对记忆的影响是什么. 这对于证明依赖于脆弱的连续吸引子概念来理解生物模拟工作记忆是基本的.

In this study, we explore perturbations and approximations of continuous attractors in the space of dynamical models.

We first report on the differences and similarities between the various structurally stable dynamics in the vicinity of continuous attractors in the space of dynamical systems models.

Our analysis reveals the presence of a "ghost" continuous attractor (a.k.a. slow manifold) in all of them (Sec. 2).

By assuming normal hyperbolicity we separate the time scales to obtain a decomposition of the dynamics by separating out the fast flow normal to and the slow flow within the slow manifold. We derive theoretical results that ensure the existence of a slow manifold and determine its closeness to a continuous attractor (Sec. 3).

在本研究中, 我们探索了动力学模型空间中连续吸引子的微扰和近似.

我们首先报告了动力系统模型空间中连续吸引子附近各种结构稳定动力学之间的差异和相似性.

我们的分析揭示了它们中都存在一个 "幽灵" 连续吸引子 (又名慢流形) (第2节).

通过假设法向 双曲, 我们 分离时间尺度, 通过分离快流 (垂直于慢流形) 和慢流 (在慢流形内) 来获得动力学的分解. 我们推导了理论结果, 确保慢流形的存在, 并确定其与连续吸引子的接近程度 (第3节).

We explore task-trained recurrent neural networks (RNNs) to show that these systems appear naturally as solutions to the task (Sec. 4) and that their generalization capabilities can easily be studied as the distance to the continuous attractor (Sec. 5).

The proposed decomposition applied to theoretical models and task-trained RNNs reveals a "universal motif" of analog memory mechanism with various potential topologies. This leads to the connection of different systems with different topologies as approximate continuous attractors (Sec. 6).

Our theory guarantees that systems close to a continuous attractor (in the space of vector fields) will have similar behavior to it, implying that the concept of continuous attractors remains a crucial framework for understanding the neural computation underlying analog memory (Sec. 3.4).

我们探索了任务训练的循环神经网络 (RNNs), 以显示这些系统自然地作为任务的解决方案出现 (第 4 节), 并且它们的 泛化能力 可以很容易地作为与连续吸引子的距离来研究 (第 5 节).

将所提出的分解应用于理论模型和任务训练的 RNNs 揭示了具有各种潜在拓扑结构的模拟记忆机制的 "通用模式". 这导致了不同拓扑结构的不同系统作为近似连续吸引子的连接 (第 6 节).

我们的理论保证了接近连续吸引子的系统 (在向量场空间中) 将具有类似的行为, 这意味着连续吸引子的概念仍然是理解模拟记忆背后的神经计算的重要框架 (第 3.4 节).

A critique of pure continuous attractors

We will first lay out a number of observations about the dynamics of bifurcations and approximations of continuous attractors used in theoretical neuroscience.

Ordinary differential equations (ODEs) are commonly used to describe the dynamical laws governing the temporal evolution of firing rates or latent population states.

In this framework, neural systems are viewed as implementing the continuous time evolution of neural states to perform computations.

We will consider a continuous attractor as a mechanism that implements analog memory computation: carrying a particular memory representation over time.

我们首先将提出一些关于理论神经科学中使用的连续吸引子的分岔和近似动力学的观察.

常微分方程 (ODEs) 通常用于描述控制 放电率潜在群体状态 时间演化的动力学规律.

在这个框架中, 神经系统被视为实现神经状态的连续时间演化以执行计算.

我们将考虑连续吸引子作为实现模拟记忆计算的机制: 含时的特定记忆表征.

To define it formally, let $\mathbf{x}(t)\in\mathbb{R}^{d}$ denote the neural state, and $\dot{\mathbf{x}}=\mathbf{f}(\mathbf{x})$ represent its dynamics.

Let $\mathcal{M}\subset \mathbb{R}^{d}$ be a manifold. We say $\mathcal{M}$ is a continuous attractor, if

(1) every state on the manifold is a fixed point, $\forall\mathbf{x}\in\mathcal{M}$, $\mathbf{f} (\mathbf{x}) = 0$, and

(2) the fixed points are marginally stable tangent to the manifold and stable normal to the manifold.

为了正式定义它, 让 $\mathbf{x}(t)\in\mathbb{R}^{d}$ 表示神经状态, 并且 $\dot{\mathbf{x}}=\mathbf{f}(\mathbf{x})$ 表示其动力学.

设流形 $\mathcal{M}\subset \mathbb{R}^{d}$. 若其满足以下条件, 我们称 $\mathcal{M}$ 为连续吸引子:

(1) 流形上的每个状态都是一个不动点, $\forall\mathbf{x}\in\mathcal{M}$, $\mathbf{f} (\mathbf{x}) = 0$, 以及

(2) 不动点在流形切线方向上是边际稳定的, 在流形法线方向上是稳定的.

In other words, the continuous attractor is a continuum of equilibrium points such that the neural state near the manifold is attracted to it, and on the manifold, the state does not move.

Marginal stability implies that continuous systems are structurally unstable, meaning that small perturbations or variations in the system's parameters lead to significant changes in the system's behavior or stability.

We will now study some examples of continuous attractors and how perturbations change their dynamics.

换句话说, 连续吸引子是一系列平衡点, 使得流形附近的神经状态被吸引到它上面, 并且在流形上, 状态不会移动.

边际稳定性意味着连续系统在结构上是不稳定的, 这意味着系统参数的小微扰或变化会导致系统行为或稳定性的显著变化.

我们现在将研究一些连续吸引子的例子, 以及微扰如何改变它们的动力学.

Motivating example: bounded line attractor

As an illustrative example, we can construct a line attractor (a continuous attractor with a line manifold) as follows:

$$ \dot{\mathbf{x}} = -\mathbf{x} + [\mathbf{Wx} + \mathbf{b}]_{+} $$

where $W = \begin{bmatrix}0 & -1 \\ -1 & 0 \end{bmatrix}$ and $b = \begin{bmatrix} 1\\ 1 \end{bmatrix}$, and $[\cdot]_{+} = \max{(0, \cdot)}$ is the threshold nonlinearity per unit.

作为一个说明性的例子, 我们可以构造一个线吸引子 (具有线流形的连续吸引子)如下:

$$ \dot{\mathbf{x}} = -\mathbf{x} + [\mathbf{Wx} + \mathbf{b}]_{+} \tag{1} $$

其中 $W = \begin{bmatrix}0 & -1 \\ -1 & 0 \end{bmatrix}$ 和 $b = \begin{bmatrix} 1\\ 1 \end{bmatrix}$, 并且 $[\cdot]_{+} = \max{(0, \cdot)}$ 是各单元的阈值非线性.

We get ̇$\dot{\mathbf{x}} = 0$ on the $x_{1} = −x_{2} + 1$ line segment in the first quadrant as the manifold (Fig. 1A, left; black line).

Linearization of the fixed points on the manifold exhibits two eigenvalues, $0$ and $−2$;

the $0$ eigenvalue allows the continuum of fixed points, while $−2$ makes the flow normal to the manifold attractive (Fig. 1A, left; flow field).

我们在第一象限的 $x_{1} = −x_{2} + 1$ 线段上得到 $\dot{\mathbf{x}} = 0$ 作为流形 (图1A, 左; 黑线).

流形上不动点的线性化显示了两个特征值, $0$ 和 $−2$;

$0$ 特征值允许不动点的连续体, 而 $−2$ 使得流形法线方向的流动具有吸引性 (图1A, 左; 流场).

Figure 1: The critical weakness of continuous attractors is their inherent brittleness as they are rare in the parameter space, i.e., infinitesimal changes in parameters destroy the continuous attractor implemented in RNNs.

Some of the structure seems to remain; there is an invariant manifold that is topologically equivalent to the original continuous attractor.

(A) Phase portraits for the bounded line attractor (Eq. (1)). Under perturbation of parameters, it bifurcates to systems without the continuous attractor.

(B) The low-rank ring attractor approximation (Sec. (S3.4)). Different topologies exist for different realizations of a low-rank attractor: different numbers of fixed points (4, 8, 12), or a limit cycle (right bottom). Yet, they all share the existence of a ring invariant set.

图 1: 连续吸引子的关键弱点是它们固有的脆弱性, 因为它们在参数空间中是罕见的, 即参数的微小变化会破坏 RNNs 中实现的连续吸引子.

一些结构似乎仍然存在; 存在一个与原始连续吸引子拓扑等价的不变流形.

(A) 有界线吸引子 (方程 (1)) 的相位图. 在参数微扰下, 它分岔为没有连续吸引子的系统.

(B) 低秩 环吸引子近似 (第 (S3.4) 节). 不同的低秩吸引子实现存在不同的拓扑结构: 不同数量的不动点 (4、8、12), 或一个极限环 (右下). 然而, 它们都共享环不变集的存在.

In general, continuous attractors are not only structurally unstable 34, they bifurcate almost certainly for an arbitrary perturbation of $\mathbf{f}$ . In this example, small changes to the parameters $(\mathbf{b}, \mathbf{W})$ perturb the eigenvalues and any to the $0$ eigenvalue destroys the continuous attractor: it bifurcates to either a single stable fixed point (Fig. 1A, top) or two stable fixed points separated with a saddle-node in between (Fig. 1A, bottom).

一般来说, 连续吸引子不仅在结构上是不稳定的, 它们几乎肯定会因为 $\mathbf{f}$ 的任意微扰而发生分岔. 在这个例子中, 对参数 $(\mathbf{b}, \mathbf{W})$ 的微小变化微扰了特征值, 任何对 $0$ 特征值的微扰都会破坏连续吸引子: 它会分岔为一个单一的稳定不动点 (图1A, 上) 或两个稳定不动点, 中间有一个鞍结点 (图1A, 下).

However, interestingly, after bifurcation, continuous attractors seemingly tend to leave a "ghost" manifold topologically equivalent to the original continuous attractor (note the slow speed). Furthermore, the flow after the bifurcation is contained in the ghost manifold, i.e., it is an invariant manifold.

This phenomenon, wherein a continuous attractor is approximated by a manifold within the neural space along which the drift occurs at a very slow pace, has previously been discussed.

然而, 有趣的是, 在分岔之后, 连续吸引子似乎倾向于留下一个与原始连续吸引子拓扑等价的 "幽灵" 流形 (注意速度很慢). 此外, 分岔后的流动包含在幽灵流形中, 即它是一个不变流形.

这个现象, 其中连续吸引子被神经空间中的一个流形近似, 在该流形上漂移以非常缓慢的速度发生, 之前已经被讨论过.

Theoretical models of ring attractors

For circular variables such as the goal-direction (e.g. for navigation and working memory for communication in bees) or head direction, the temporal integration, and working memory functions are naturally solved by a ring attractor (continuous attractor with a ring topology).

Other examples include integration of evidence for continuous perceptual judgments, e.g. a hunting predator that needs to compute the net direction of motion of a large group of prey.

对于圆形变量, 例如目标方向 (例如用于导航和蜜蜂通信的工作记忆) 或头部方向, 时间积分和工作记忆功能自然地由环吸引子 (具有环拓扑的连续吸引子)解决.

其他例子包括对连续感知判断的 证据积分, 例如需要计算大量猎物群体的净运动方向的狩猎捕食者.

In this section, we investigate the bifurcations of various implementations of continuous attractors. Continuous attractor network models of the head direction are based on the interactions of neurons that (anatomically) form a ring-like overlapping neighbor connectivity.

Similarly to the line attractor, the ring attractor bifurcates with almost any perturbation to the network dynamics. However, the resulting dynamics continue to follow a familiar pattern: they remain confined to a ghost manifold that closely approximates the original continuous attractor.

在本节中, 我们研究了连续吸引子各种实现的分岔. 头部方向的连续吸引子网络模型基于神经元之间的相互作用, 这些神经元 (解剖上) 形成类似环状的重叠近邻连接.

与线吸引子类似, 环吸引子在网络动力学的几乎任何微扰下都会发生分岔. 然而, 结果动力学继续遵循熟悉的模式: 它们仍然被限制在一个幽灵流形中, 该流形紧密地近似原始连续吸引子.

Piecewise-linear ring attractor model of the central complex

Firstly, we discuss perturbations of a continuous ring attractor recently proposed as a model for the head direction representation in fruit flies.

This model is composed of $N$ heading-tuned neurons with preferred headings $\theta_{j}\in\{2\pi i/N\}_{i=1\cdots N}$ radians (Sec. S3.2).

For sufficiently strong local excitation (given by the parameter $J_{E}$) and broad inhibition ($J_{I}$), this network will generate a stable bump of activity corresponding to the head direction. This continuum of fixed points forms a $N$-sided polygon.

首先, 我们讨论最近提出的作为果蝇头部方向表示模型的连续环吸引子的微扰.

该模型由 $N$ 个具有偏好方向 $\theta_{j}\in\left\{\frac{2\pi i}{N}\right\}_{i=1\cdots N}$ 弧度的头部调谐神经元组成 (第 S3.2 节).

对于足够强的 局部激励 (由参数 $J_{E}$ 给出) 和 广域抑制 ($J_{I}$), 该网络将生成一个稳定的活动峰, 对应于头部方向. 这个不动点的连续体形成了一个 $N$ 边形.

We evaluate the effect of parametric perturbations of the form $\mathbf{W}\leftarrow \mathbf{W} + \mathbf{V}$ with $\mathbf{V}_{i,j}\overset{iid}{\sim}\mathcal{N}(0, 0.01)$ on a network of size $N = 6$ (forming an hexagon, see also Sec. S3).

We found that the continuum of fixed points can collapse to between $2$ and $12$ isolated fixed points (Fig. 2A).

As far as we know, this bifurcation from a ring of equilibria to a saddle and node has not been described previously in the literature. The probability of each type of bifurcation was numerically estimated (Sec. S3.2).

Surprisingly, the number of fixed points is maintained throughout a range of perturbation sizes and hence depends only on the direction of the perturbation (Fig. S8).

我们评估了在大小为 $N = 6$ 的网络上, 形式为 $\mathbf{W}\leftarrow \mathbf{W} + \mathbf{V}$ 的参数微扰的影响, 其中 $\mathbf{V}_{i,j}\overset{iid}{\sim}\mathcal{N}(0, 0.01)$ (形成一个六边形, 另见第 S3 节).

我们发现, 不动点的连续体可以坍缩为 $2$ 到 $12$ 个孤立的不动点 (图 2A).

据我们所知, 从平衡环到鞍点和节点的这种分岔以前在文献中没有描述过. 每种类型分岔的概率是通过数值估计的 (第 S3.2 节).

令人惊讶的是, 不动点的数量在一系列微扰大小中保持不变, 因此仅取决于微扰的方向 (图 S8).

Bump attractor model

A well-established approach to form a ring attractor in the limit of large network is with a connection matrix $W$ with entries that follow a circular Gaussian function of $i − j$ (Sec. S3.3). This type of ring attractor network can support a stable "activity bump" that can move around the ring of nonlinear neurons in correspondence with changes in head direction.

For finite-sized networks, the dynamics are constrained to an attractive invariant ring, covered with $N$ stable fixed points for a network of size $N$ (Fig. 2B). For such networks the number of fixed points can change with the size of the perturbation (Fig. S8).

一个在大网络极限下形成环吸引子的成熟方法是使用连接矩阵 $W$, 其条目遵循 $i − j$ 的 圆高斯函数 (第 S3.3 节). 这种类型的环吸引子网络可以支持一个稳定的 "活动峰", 它可以围绕非线性神经元的环移动, 以对应于头部方向的变化.

对于有限大小的网络, 动力学被约束在一个有吸引力的不变环上, 对于大小为 $N$ 的网络, 覆盖有 $N$ 个稳定不动点 (图 2B). 对于这样的网络, 不动点的数量可以随着微扰的大小而变化 (图 S8).

Perturbations of different implementations and approximations of ring attractors lead to bifurcations that all leave the ring invariant manifold intact.

For each model, the network dynamics is constrained to a ring manifold with stable fixed points (green) and saddle nodes (red).

(A) Perturbations to Noorman et al.. The ring attractor can be perturbed in systems with an even number of fixed points (FPs) up to $2N$ (stable and saddle points are paired).

(B) Perturbations to a $\tanh$ approximation of a ring attractor Seeholzer et al..

(C) Different Embedding Manifolds with Population-level Jacobians (EMPJ) approximations of a ring attractor.

微扰不同实现和近似的环吸引子会导致分岔, 这些分岔都保持了环不变流形的完整性.

对于每个模型, 网络动力学被约束在一个具有稳定不动点 (绿色) 和鞍结点 (红色) 的环流形上.

(A) 对 Noorman 等人的微扰. 环吸引子可以在具有偶数个不动点 (FPs) 最多 $2N$ 的系统中被微扰 (稳定点和鞍点成对).

(B) 对 Seeholzer 等人的环吸引子的 $\tanh$ 近似的微扰.

(C) 具有群体级 Jacobian 矩阵 (EMPJ) 的不同嵌入流形的环吸引子近似.

Low-rank ring attractor model

Low-rank networks can be used to approximate ring attractors. In the limit of infinite-size networks, one can construct a ring attractor through a rank 2 network by constraining the overlap of the right- and left-connectivity vectors (see Sec. S3.4).

However, in simulations of finite-size networks with this constraint, the dynamics instead always converge to a small number of stable fixed points arranged on a ring (Fig. 1B).

低秩网络可用于近似环吸引子. 在无限大小网络的极限下, 可以通过约束右连接向量和左连接向量的重叠来构造一个秩为 2 的环吸引子网络 (见第 S3.4 节).

然而, 在具有此约束的有限大小网络的模拟中, 动力学总是收敛到排列在环上的少量稳定不动点 (图 1B).

Embedding Manifolds with Population-level Jacobians

Approximate ring attractors can be constructed by constraining the connectivity so that the networks Jacobian satisfies certain requirements for a ring attractor to exist(see Sec. S3.5).

The models constructed with this method also can contain an invariant ring manifold on which the dynamics contain stable and saddle fixed points (Fig. 2C).

It has been observed that approximate continuous attractor emerge in networks trained on sampled points with other methods as well.

近似环吸引子可以通过约束连接来构建, 使得网络的 Jacobian 矩阵满足存在环吸引子的某些要求 (见第 S3.5 节).

用这种方法构建的模型也可以包含一个不变的环流形, 其上的动力学包含稳定和鞍不动点 (图 2C).

已经观察到, 近似连续吸引子也会出现在使用其他方法在采样点上训练的网络中.

Similarity between all bifurcations and approximations of continuous attractors

In all discussed models of ring attractors, we verify that they suffer from the fine-tuning problem.

However, importantly, we also observe in all the systems the existence of ghosts of the continuous attractor (either through bifurcation or from finite-size effects) in the form of an attractive invariant manifold.

Therefore, while they are not strictly a continuous attractor in the mathematical sense, they are approximate ring attractors in the sense that the fixed points and connecting orbits still form a circle.

在所有讨论的环吸引子模型中, 我们确认它们全都存在微调问题.

然而, 重要的是, 我们还观察到在所有系统中都存在连续吸引子的幽灵 (无论是通过分岔还是有限大小效应), 以具有吸引力的不变流形的形式.

因此, 虽然它们在数学意义上不严格是连续吸引子, 但它们在不动点和连接轨道仍然形成一个圆的意义上是近似环吸引子.

Is this lawful degradation a universal phenomenon? And if so, how does it relate to the size of the perturbation? And what are the implications for the memory performance of these approximations? (Sec. 3).

Do these approximations appear as natural solutions to the memory storage problem? (Sec. 4).

And if so, how well do they generalize to longer time requirements? (Sec. 5).

Finally, are continuous attractors in practice still useful as an idealized model of how animals represent continuous variables? (Sec. 6).

这种合法的退化是普遍现象吗? 如果是, 它与微扰的大小有何关系? 这些近似对记忆性能有什么影响? (第3节).

这些近似是否作为记忆存储问题的自然解决方案出现? (第4节).

如果是, 它们在更长时间要求下的泛化能力如何? (第5节).

最后, 作为动物如何表征连续变量的理想化模型, 连续吸引子在实践中仍然有用吗? (第6节).

Theory of Approximate Continuous Attractors

In this section, we theoretically answer in an implementation-agnostic manner the degradation questions posed from the exploration. To do so, we apply invariant manifold theory to continuous attractor models and translate the results for the neuroscience audience (see also Sec. S1).

在本节中, 我们以与实现无关的方式从理论上回答了探索中提出的 退化问题. 为此, 我们将不变流形理论应用于连续吸引子模型, 并将结果转译给神经科学受众 (另见第 S1 节).

Persistent Manifold Theorem

First, we argue that the lawful degradation into a system with a slow manifold is universally guaranteed (as long as the perturbation is small, and the continuous attractor was normally hyperbolic).

Let $l$ be the intrinsic dimension of the manifold of equilibria that defines the continuous attractor. Given a perturbation $\mathbf{p}(\mathbf{x})$ to the ODE that induces a bifurcation,

$$ \dot{\mathbf{x}} = \mathbf{f}(\mathbf{x}) + \epsilon\mathbf{p}(\mathbf{x}) $$

where $||\mathbf{p}(\cdot)||_{\infty} = 1$ and $\epsilon > 0$ is the bifurcation parameter, we can reparameterize the dynamics around the manifold with coordinates $\mathbf{y}\in\mathbb{R}^{l}$ and the remaining ambient space with $\mathbf{z}\in \mathbb{R}^{d−l}$.

首先, 我们认为退化为具有慢流形的系统是普遍保证的 (只要微扰很小, 并且连续吸引子是法向双曲的). 设 $l$ 为定义连续吸引子的平衡流形的 内在维数. 给定对 ODE 的微扰 $\mathbf{p}(\mathbf{x})$, 它会引起分岔,

$$ \dot{\mathbf{x}} = \mathbf{f}(\mathbf{x}) + \epsilon\mathbf{p}(\mathbf{x})\tag{2} $$

其中 $||\mathbf{p}(\cdot)||_{\infty} = 1$ 且 $\epsilon > 0$ 是分岔参数, 我们可以使用坐标 $\mathbf{y}\in\mathbb{R}^{l}$ 对流形附近的动力学进行重新参数化, 并使用 $\mathbf{z}\in \mathbb{R}^{d−l}$ 对剩余的环境空间进行参数化.

To describe an arbitrary bifurcation of interest, we introduce a sufficiently smooth function $\mathbf{g}$ and $\mathbf{h}$, such that the following system is equivalent to the original ODE:

$$ \dot{\mathbf{y}} = \epsilon&\mathbf{g}(\mathbf{y},\mathbf{z}, \epsilon)\quad (\mathrm{tangent})\\ \dot{\mathbf{z}} = &\mathbf{h}(\mathbf{y},\mathbf{z}, \epsilon)\quad (\mathrm{normal}) $$

where $\epsilon = 0$ gives the condition for the continuous attractor $\dot{\mathbf{y}} = \mathbf{0}$. We denote the corresponding manifold of $l$ dimensions as $\mathcal{M}_{0}:=\{(\mathbf{y}, \mathbf{z})|\mathbf{h}(\mathbf{y},\mathbf{z}, 0) = \mathbf{0}\}$.

为了描述感兴趣的任意分岔, 我们引入一个足够光滑的函数 $\mathbf{g}$ 和 $\mathbf{h}$, 使得以下系统与原始 ODE 等价:

$$ \begin{align} \dot{\mathbf{y}} = \epsilon&\mathbf{g}(\mathbf{y},\mathbf{z}, \epsilon)\quad (\mathrm{tangent})\tag{3}\\ \dot{\mathbf{z}} = &\mathbf{h}(\mathbf{y},\mathbf{z}, \epsilon)\quad (\mathrm{normal})\tag{4} \end{align} $$

其中 $\epsilon = 0$ 给出了连续吸引子的条件 $\dot{\mathbf{y}} = \mathbf{0}$. 我们将相应的 $l$ 维流形表示为 $\mathcal{M}_{0}:=\{(\mathbf{y}, \mathbf{z})|\mathbf{h}(\mathbf{y},\mathbf{z}, 0) = \mathbf{0}\}$.

We say the flow around the manifold is normally hyperbolic, if the flow normal to the manifold is hyperbolic, i.e. the Jacobians $\nabla_{\mathbf{z}}\mathbf{h}$ evaluated at any point on the $\mathcal{M}_{0}$ has $d − l$ eigenvalues with their real part uniformly bounded away from zero, and $\nabla_{\mathbf{y}}\mathbf{g}$ has $l$ eigenvalues with zero real part.

More specifically, for continuous attractors, the real parts of the eigenvalues of $\nabla_{\mathbf{z}}\mathbf{h}$ are strictly negative, representing sufficiently strong attractive flow toward the manifold.

我们称流形附近的流是 法向双曲 的, 如果流形法向的流动是双曲性的, 即在 $\mathcal{M}_{0}$ 上的任意点处评估的 Jacobian 矩阵 $\nabla_{\mathbf{z}}\mathbf{h}$ 具有 $d − l$ 个特征值, 并且这些特征值的实部都一致远离零, 并且 $\nabla_{\mathbf{y}}\mathbf{g}$ 具有 $l$ 个特征值, 其实部为零.

更具体地, 对于连续吸引子, $\nabla_{\mathbf{z}}\mathbf{h}$ 的特征值的实部严格为负, 表示足够强的吸引流向流形.

Equivalently, for the ODE, $\dot{\mathbf{x}} = \mathbf{f}(\mathbf{x})$, the variational system is of constant rank and has exactly $(d − l)$ eigenvalues with negative real parts uniformly away from zero and $l$ eigenvalues with zero real parts everywhere along the continuous attractor.

等价地, 对于 ODE $\dot{\mathbf{x}} = \mathbf{f}(\mathbf{x})$, 变分系统具有常数秩, 并且在连续吸引子上处处具有恰好 $(d − l)$ 个特征值的实部为负且一致远离零, 并且具有 $l$ 个特征值的实部为零.

Figure 3: Persistent manifold theorem applied to compact continuous attractor guarantees the flow on the slow manifold $\mathcal{M}_{\epsilon}$ is invariant and continues to be attractive. The dashed line is a trajectory "trapped" in the slow manifold (which has the same topology as the continuous attractor).

图 3: 持久流形定理应用于紧致连续吸引子, 确保慢流形 $\mathcal{M}_{\epsilon}$ 上的流动是不变的, 并且继续具有吸引性. 虚线是被 "束缚" 在慢流形中的轨迹 (其拓扑结构与连续吸引子相同).

For any parameterization $\mathbf{g}$, $\epsilon > 0$ induces a bifurcation of the continuous attractor. What can we say about the fate of the perturbed system?

The continuous dependence theorem says that the trajectories will change continuously as a function of $\epsilon$ without a guarantee on how quickly they change.

However, the topological structure and the asymptotic behavior of trajectories change discontinuously due to the bifurcation.

Yet, surprisingly, there is a strong connection in the geometry due to Fenichel's theorem. We informally present a special case due to Jones:

Theorem 1 (Persistent Manifold). Let $\mathcal{M}_{0}$ be a connected, compact, normally hyperbolic manifold of equilibria originating from a sufficiently smooth ODE. For a sufficiently small perturbation $\epsilon > 0$, there exists a manifold $\mathcal{M}_{\epsilon}$ diffeomorphic to $\mathcal{M}_{0}$ and invariant under the flow of Eq.(3)-(4).

对于任何参数化 $\mathbf{g}$, $\epsilon > 0$ 都会引起连续吸引子的分岔. 我们能说些什么关于微扰系统的命运?

连续依赖定理 表明, 轨迹将作为 $\epsilon$ 的函数连续变化, 但不能保证它们变化的速度.

然而, 由于分岔, 轨迹的拓扑结构和渐近行为会不连续地变化.

令人惊讶的是, 由于 Fenichel 定理, 在几何上存在强烈的联系. 我们非正式地提出一个由 Jones 提出的特殊情况:

定理 1 (持久流形). 设 $\mathcal{M}_{0}$ 为一个连通的、紧致的、法向双曲的平衡流形, 源自一个足够光滑的 ODE. 对于足够小的微扰 $\epsilon > 0$, 存在一个流形 $\mathcal{M}_{\epsilon}$, 它与 $\mathcal{M}_{0}$ 微分同胚, 并且在方程 (3)-(4) 的流动下是不变的.

The manifold $\mathcal{M}_{\epsilon}$ is called the slow manifold which is no longer necessarily a continuum of equilibria.

However, the invariance implies that trajectories remain within the manifold except potentially at the boundary.

Furthermore, the non-zero flow on the slow manifold is slow and given in the $\epsilon\to 0$ limit as $\frac{\mathrm{d}\mathbf{y}}{\mathrm{d}\tau} = \mathbf{g}(c^{\epsilon}(\mathbf{y}), \mathbf{y}, 0)$ where $\tau = \epsilon t$ is a rescaled time and $c\epsilon(\cdot)$ parameterizes the $l$ dimensional slow manifold.

In addition, the stable manifold of $\mathcal{M}_{0}$ is similarly persistent, implying that the manifold $\mathcal{M}_{\epsilon}$ remains attractive.

Finally, the persisting invariant manifold is very close in space to the original continuous attractor (see also Theorem 3).

流形 $\mathcal{M}_{\epsilon}$ 被称为 慢流形, 它不再必定是平衡的连续体.

然而, 不变性意味着轨迹保持在流形内, 除了可能在边界处.

此外, 慢流形上的非零流动是缓慢的, 并且在 $\epsilon\to 0$ 极限下给出为 $\frac{\mathrm{d}\mathbf{y}}{\mathrm{d}\tau} = \mathbf{g}(c^{\epsilon}(\mathbf{y}), \mathbf{y}, 0)$, 其中 $\tau = \epsilon t$ 是重新缩放的时间, $c^{\epsilon}(\cdot)$ 参数化了 $l$ 维慢流形.

此外, $\mathcal{M}_{0}$ 的稳定流形同样是持续的, 这意味着流形 $\mathcal{M}_{\epsilon}$ 保持具有吸引力.

最后, 持续存在的不变流形在空间上非常接近原始连续吸引子 (另见定理 3).

These conditions are met for the examples in Fig. 1 (see Sec. S2.1 for the corresponding fast-slow reparametrization).

As the theory predicts, it bifurcates into a 1-dimensional slow manifold (Fig. 1, dark-colored regions) that contains fixed points and connecting orbits and is overall still attractive.

Furthermore, Fenichel's Persistent Manifold theorem explains the bifurcation structure of the theoretical models discussed in Sec. 2.2. Because continuous ring attractors are bounded, they persist as an invariant manifold and remain attractive under small perturbations.

这些条件满足图 1 中的示例 (有关相应的快慢重新参数化, 请参见第 S2.1 节).

正如理论所预测的, 它分岔为一个 1 维慢流形 (图 1, 深色区域), 其中包含不动点和连接轨道, 并且总体上仍然具有吸引力.

此外, Fenichel 的持久流形定理解释了第 2.2 节中讨论的理论模型的分岔结构. 由于连续环吸引子是有界的, 它们在小微扰下作为不变流形持续存在并保持具有吸引力.

Fast-slow decomposition and the revival of continuous attractors

Consider a behaviorally relevant timescale for working memory, for example, roughly up to a few tens of seconds.

If the relevant dynamical system is orders of magnitude slower, for example, $1000$ sec or longer, its effect is too slow to have a practical impact on the behavior. This clear gap in the fast and slow time scales can be recast as normal hyperbolicity of the slow manifold by relaxing the zero real part to a separation of time scales (reciprocal of eigenvalues or Lyapunov exponents).

In other words, the attractive flow normal to the manifold needs to be uniformly faster than the flow on the slow manifold. By taking the limit of the slow flow on the manifold to arbitrarily long time constant (i.e., to zero flow), we achieve the reversal of the persistent manifold theorem.

考虑一个与行为相关的工作记忆时间尺度, 例如, 大约几秒到几十秒.

如果相关的动力系统慢几个数量级, 例如, $1000$ 秒或更长, 它的影响太慢而无法对行为产生实际影响. 快慢时间尺度之间的这种明显差距可以通过将零实部宽松为 时间尺度分离 (特征值或 Lyapunov 指数的倒数) 以重新刻画慢流形的法向双曲性.

换句话说, 流形法向的吸引流动需要比慢流形上的流动均匀地更快. 通过将流形上慢流动的极限取为任意长的时间常数 (即, 零流动), 我们实现了持久流形定理的逆转.

Proposition 1 (Revival of continuous attractor). Let $\mathcal{M}_{\epsilon}$ be a connected, compact, attractive, normally hyperbolic slow manifold (as parametrized by Eq. (3)-(4)). Let the uniform norm of the flow tangent to the manifold be $||\dot{\mathbf{y}}||_{\infty} = \eta$. There exists a perturbation with uniform norm at most $\eta$ that induces a bifurcation to a continuous attractor manifold.

命题 1 (为连续吸引子招魂). 设 $\mathcal{M}_{\epsilon}$ 为一个连通的、紧致的、有吸引力的、法向双曲的慢流形 (如方程 (3)-(4) 参数化). 设流形切线方向的流的均匀范数为 $||\dot{\mathbf{y}}||_{\infty} = \eta$. 存在一个均匀范数至多为 $\eta$ 的微扰, 它会引起分岔到一个连续吸引子流形.

An explicit perturbation is derived in Sec. S5.1. This makes the uniform norm of the vector field on a (slow) manifold a useful measure of the distance of an approximation to a continuous attractor. Prop. 1 can be extended to the case where the invariant manifold has additional dynamics to which the output mapping is invariant (see Theorem 7). These systems can be perturbed onto a decomposable system where one of the subsystems has a slow flow.

一个明确的微扰在第 S5.1 节中导出. 这使得 (慢) 流形上的向量场的均匀范数成为衡量 某近似到连续吸引子的距离 的有用指标. 命题 1 可以扩展到不变流形具有额外动力学的情况, 对于这些动力学, 输出映射是不变的 (见定理 7). 这些系统可以被微扰到一个可分解系统, 其中一个子系统具有慢流动.

Relevance of dynamics on the memory performance of the slow manifold

Third, we relate the flow of the manifold (and, through Prop. 1, the size of the perturbation) to the memory error of the approximation in short-time scale. We also discuss the implications of the theoretical insights on the memory error in the asymptotic time scale.

第三, 我们将流形上的流 (并通过命题 1, 微扰的大小) 与短时间尺度下近似的记忆误差联系起来. 我们还讨论了理论见解对渐近时间尺度下记忆误差的影响.

In the short-time scale the memory performance is bounded by the uniform norm of the flow tangent to the manifold. Let $x_{0}\in \mathcal{M}$, and $\varphi = \mathbf{p}(\cdot)|_{\mathcal{M}}$ be the vector field restricted to the manifold (following the notation in Eq. 2).

在短时间尺度下, 记忆性能受限于流形切线方向的流的均匀范数.

设 $x_{0}\in \mathcal{M}$, 并且 $\varphi = \mathbf{p}(\cdot)|_{\mathcal{M}}$ 是限制在流形上的向量场 (遵循方程 2 中的符号).

The average deviation from the initial memory $\mathbf{x}_{0}$ over time is bounded linearly (derivation in Sec. S6):

$$ \frac{1}{\mathrm{vol}\,\mathcal{M}} \int_{\mathcal{M}} |\mathbf{x}(t,\mathbf{x}_{0}) - \mathbf{x}_{0}| \mathrm{d}\mathbf{x}_{0} \leq t||\varphi||_{\infty},\quad (\text{error bound}) $$

Note that this bound is the worst case and tighter for sufficiently small $t\geq 0$.

Furthermore, for compact invariant manifolds the error is bounded by the diameter of the manifold and hence this bound becomes irrelevant for $t$ large.

初始记忆 $\mathbf{x}_{0}$ 的含时平均偏差被线性限制 (推导见第 S6 节):

$$ \frac{1}{\mathrm{vol}\,\mathcal{M}} \int_{\mathcal{M}} |\mathbf{x}(t,\mathbf{x}_{0}) - \mathbf{x}_{0}| \mathrm{d}\mathbf{x}_{0} \leq t||\varphi||_{\infty},\quad (\text{error bound}) $$

注意, 这个界限是最坏情况, 并且对于足够小的 $t\geq 0$ 更紧.

此外, 对于紧致的不变流形, 误差受流形直径的限制, 因此对于大的 $t$, 这个界限变得无关紧要.

While the uniform norm gives insight on the short-time scale behavior of the perturbed ODE, we expect that working memory tasks generalize to longer durations.

The long-time scale behavior on the slow manifold is dominated by the stability structure, i.e., the topology of the dynamics. Although we have seen numerous topologies in Sec. 2, the Persistent Manifold Theorem says that this variability is fundamentally limited, especially in low dimensions (see for more details Sec. S4).

虽然均匀范数提供了对微扰 ODE 的短时间尺度行为的洞察, 但我们期望工作记忆任务可以泛化到更长的持续时间.

慢流形上的长时间尺度行为由稳定性结构主导, 即动力学的拓扑结构. 尽管我们在第 2 节中看到了许多拓扑结构, 但持久流形定理表明, 这种变异性在根本上是有限的, 尤其是在低维度下 (有关更多详细信息, 请参见第 S4 节).

This is especially relevant as previous works have identified a low-dimensional organization of neural activity to explain the brain's ability to adapt behavioral responses to changing stimuli and environments.

For a ring attractor, this implies that the stability structure of the invariant manifold is either

(1) composed of an equal number of stable fixed points and saddle nodes, placed alternatingly and with connecting orbits, or

(2) a limit cycle.

These different stability structures have different generalization properties (see Sec. 5). In more complex scenarios, such as two-dimensional attractors, fixed points can coexist with limit cycles, creating a rich tapestry of possible attractors.

这尤其相关, 因为以前的研究已经确定了神经活动的低维组织, 以解释大脑适应行为反应以应对不断变化的刺激和环境的能力.

对于环吸引子, 这意味着不变流形的稳定性结构要么

(1) 由相等数量的稳定不动点和鞍结点组成, 交替放置并具有连接轨道, 要么

(2) 是一个极限环.

这些不同的稳定性结构具有不同的泛化特性 (见第 5 节). 在更复杂的场景中, 例如二维吸引子, 不动点可以与极限环共存, 从而创造出丰富的可能吸引子.

Implications on experimental neuroscience

Animal behavior exhibits strong resilience to changes in their neural dynamics, such as the continuous fluctuations in the synapses or slight variations in neuromodulator levels or temperature. Hence, any theoretical model of neural or cognitive function that requires fine-tuning, such as the continuous attractor model for analog working memory, raises concerns, as they are seemingly biologically irrelevant.

动物行为表现出对神经动力学变化的强大弹性, 例如突触的连续波动或神经调节剂水平或温度的轻微变化. 因此, 任何需要微调的神经或认知功能理论模型, 例如用于模拟工作记忆的连续吸引子模型, 都引发了担忧, 因为它们似乎在生物学上无关紧要.

This challenge is further compounded by the structural constraints imposed by the connectome, which defines the network's architecture and limits the possible configurations of synaptic and circuit dynamics.

Moreover, unbiased data-driven models of time series data and task-trained recurrent network models cannot recover such continuous attractor theories precisely.

Our theory shows that this apparent fragility is not as devastating as previously thought: despite the "qualitative differences" in the phase portrait, the "effective behavior" of the system can be arbitrarily close, especially in the behaviorally relevant time scales.

We show that as long as the attractive flow to the memory representation manifold is fast and the flow on the manifold is sufficiently slow, it represents an approximate continuous attractor.

连接组 施加的结构约束进一步加剧了这一挑战, 它定义了网络的架构并限制了突触和回路动力学的可能配置.

此外, 无偏的 数据驱动时间序列数据模型任务训练的循环网络模型 无法精确地恢复这种连续吸引子理论.

我们的理论表明, 这种明显的脆弱性并不像以前认为的那样具有破坏性: 尽管相位肖像中存在 "定性差异", 但系统的 "有效行为" 可以任意接近, 尤其是在与行为相关的时间尺度上.

我们表明, 只要对记忆表示流形的吸引流动快速, 并且流形上的流动足够慢, 它就代表一个近似连续吸引子.

Furthermore, our theory bounds the error in working memory incurred over time for such approximate continuous attractors.

Therefore, the concept of continuous attractors remains a crucial framework for understanding the neural computation underlying analog memory, even if the ideal continuous attractor is never observed in practice.

Experimental observations that indicate the slowly changing population representations during the "delay periods" where working memory is presumably required, do not necessarily contradict the continuous attractor hypothesis.

Perturbative experiments can further measure the attractive nature of the manifold and their causal role through manipulating the memory content.

此外, 我们的理论界定了这种近似连续吸引子在工作记忆中随时间产生的误差.

因此, 即使在实践中从未观察到理想的连续吸引子, 连续吸引子的概念仍然是理解模拟记忆背后的神经计算的重要框架.

在 "延迟期" 期间, 工作记忆可能需要时, 实验观察表明群体表征缓慢变化, 并不一定与连续吸引子假说相矛盾.

微扰实验可以进一步测量流形的吸引性质及其通过操纵记忆内容所起的因果作用.

Numerical Experiments on Task-optimized Recurrent Networks

While our theory describes the abundance of approximate continuous attractors in the vicinity of a continuous attractor, it does not imply that there are no approximate solutions away from continuous attractors.

In this section, we use task-optimized RNNs as a means to search for plausible solutions for analog memory for a circular variable. We train a diverse set of RNNs, and then identify the solution type of trained RNNs to gain insights into its performance, error patterns, generalization capabilities, and, ultimately, proximity to a continuous attractor.

虽然我们的理论描述了连续吸引子附近近似连续吸引子的丰富性, 但它并不意味着在连续吸引子之外没有近似解.

在本节中, 我们使用 任务优化的 RNN 作为搜索圆形变量模拟记忆的合理解的方法. 我们训练了一组多样化的 RNN, 然后识别训练 RNN 的解类型, 以获得对其性能、误差模式、泛化能力以及最终与连续吸引子的接近性的见解.

Understanding the implemented computation in neural systems in terms of dynamical systems is a well-established practice. Researchers have analyzed task-optimized RNNs through nonlinear dynamical systems analysis and to compare those artificial networks to biological circuits.

Previously, systematic analysis of the variability in network dynamics has been surveyed in vanilla RNNs, and variations in dynamical solutions over architecture and nonlinearity have been quantified.

理解神经系统中实现的计算在动力系统方面是一种成熟的实践. 研究者已经通过非线性动力系统分析分析了 任务优化的 RNN, 并将这些人工网络与生物回路进行比较.

此前, 已经对 vanilla RNN 中网络动力学的变异性进行了系统分析, 并量化了架构和非线性上的动力学解的变化.

Furthermore, working memory mechanisms in RNNs had tendencies to find sequential or persistent representations through training depending on the task specification.

We therefore investigated to what extent training RNNs on a task uniquely determines the low-dimensional dynamics, independent of neural architectures.

We see that all the solutions have a slow invariant manifold, making all of them an instantiation of approximate continuous attractors.

此外, RNN 中的工作记忆机制在训练过程中有倾向于根据任务规范找到序列或持久表征.

因此, 我们调查了在任务上训练 RNN 在多大程度上独立于神经架构唯一地确定低维动力学.

我们发现所有的解都有一个慢不变流形, 使它们都成为近似连续吸引子的实例.

Model Architectures and Training Procedure

Building upon prior work, which has shown their capabilities on such tasks, we trained RNNs to either

(1) estimate head direction through integration of angular velocity (Fig. 4A1) or

(2) perform a memory-guided saccade task for targets on a circle (Fig. 4A2, details in Sec. S7.1 and see Sec. S7.3 for how RNNs relate to Eq. 2).

在之前工作的基础上, 已经显示了它们在此类任务中的能力, 我们训练 RNN 来

(1) 通过积分角速度来估计头部方向 (图 4A1) 或

(2) 执行圆形目标的记忆引导扫视任务 (图 4A2, 详细信息见第 S7.1 节, 并参见第 S7.3 节了解 RNN 如何与方程 2 相关).

We numerically minimized the mean squared error loss LMSE between the network output and the target output.

For each activation function and each network architecture (vanilla RNN with ReLU, tanh, and rectified tanh activation functions, LSTM, and GRU), we trained 10 networks per hidden size: 64, 128, and 256 with (hidden) state noise.

我们数值最小化网络输出与目标输出之间的均方误差损失 LMSE.

对于每个激活函数和每个网络架构 (具有 ReLU、tanh 和修正 tanh 激活函数的 vanilla RNN、LSTM 和 GRU), 我们为每个隐藏层大小训练了 10 个网络: $64$、$128$ 和 $256$, 并添加了 (隐藏) 状态噪声.

Numerical Fast-Slow Decomposition

For each trained network, we find the slow manifold by integrating the autonomous dynamics, then selecting the parts of the trajectories that have speed slower than a threshold (Sec. S7.9.1).

We identify the points on the invariant manifold from the simulated trajectories that are projected closest to a set of points in the output space relevant to the task after convergence, i.e. on the target ring.

We parametrize the one-dimensional invariant manifold by fitting a cubic spline with periodic boundary constraints to these points (black line in Fig. 4B & C). Normal hyperbolicity is measured by a gap in the timescales of the system (the eigenvalue spectrum of the linearization along points on the invariant manifold, Fig. 4E and F).

对于每个训练的网络, 我们通过积分自治动力学找到慢流形, 然后选择速度低于阈值的轨迹部分 (第 S7.9.1 节).

我们从模拟轨迹中识别不变流形上的点, 这些点在收敛后投影到与任务相关的输出空间中的一组点上, 即在目标环上.

我们通过对这些点拟合具有周期边界约束的三次样条来参数化一维不变流形 (图 4B 和 C 中的黑线). 法向双曲通过系统时间尺度的间隙来测量 (沿不变流形上的点线性化的特征值谱, 图 4E 和 F).

We find the fixed points on the invariant ring by identifying regions where the direction of the flow flips (Sec. S7.6.3). Stable fixed points are identified where the flow directions are both pointing towards this flip point, while saddle nodes are identified where they are pointing away (Fig. 4B & C.)

我们通过识别流动方向翻转的区域来找到不变环上的不动点 (第 S7.6.3 节). 稳定不动点被识别为流动方向都指向该翻转点, 而鞍结点被识别为流动方向指向远离该点 (图 4B 和 C).

Figure 4: Slow manifold approximation of different trained networks on the memory-guided saccade and angular velocity integration tasks. (A1) Output of an example trajectory on the angular velocity integration task. (A2) Output of example trajectories on the memory-guided saccade task. (B) An example fixed-point type solution to the memory-guided saccade task. Circles indicate fixed points of the system (filled for stable, empty for saddle) and the decoded angular value on the output ring is indicated with the color according to A1. (C) An example of a found solution to the angular velocity integration task. (D) An example slow-torus type solution to the memory-guided saccade task. The colored curves indicate stable limit cycles of the system. (E+F) The eigenvalue spectra for the trained networks in B and C show a gap between the first two largest eigenvalues.

图 4: 记忆引导扫视和角速度积分任务中不同训练网络的慢流形近似. (A1) 角速度积分任务中示例轨迹的输出. (A2) 记忆引导扫视任务中示例轨迹的输出. (B) 记忆引导扫视任务的不动点类型解的示例. 圆圈表示系统的不动点 (填充表示稳定, 空心表示鞍点), 输出环上的解码角值用 A1 中的颜色表示. (C) 角速度积分任务中找到的解的示例. (D) 记忆引导扫视任务中的慢环面类型解的示例. 彩色曲线表示系统的稳定极限环. (E+F) B 和 C 中训练网络的特征值谱显示前两个最大特征值之间存在间隙.

Variations in the Topologies of the Slow Manifold Solutions

To understand what solutions the RNNs found to solve the task, we investigate their memory mechanism.

For this, we dissect the dynamics of RNNs by segregating time scales to delineate the rapid flow normal to the slow manifold, and the flow on the manifold (Sec. S7.6.3).

为了理解 RNN 为解决任务找到的解, 我们调查了它们的记忆机制.

为此, 我们通过分离时间尺度来剖析 RNN 的动力学, 以划定慢流形法线方向的快速流动和流形上的流动 (第 S7.6.3 节).

All solutions involve a slow manifold with the same topology as the relevant variable in the task. The different solutions are different in their asymptotic dynamics (Fig. 4). The most often found solution is of the type fixed point ring manifold (Fig. 4B and C). These solutions are consistent with observations that persistent activity relies on discrete attractors 102,103. Less commonly found topologies includes the slow torus around a repulsive ring invariant manifold (Fig. 4D). This solution in turn is consistent with both observations of the possibility of using non-constant dynamics for memory storage 19,104 and neuronal circuits underlying persistent representations despite time-varying activity 105. All stability structures (fixed points and limit cycles) are mapped close to the target output circle (Figs. S15, S19, S20). We also verify the normal hyperbolicity of the trained networks shown in Fig. 4B and C. The largest eigenvalue of the Jacobian fluctuates around zero (the invariant manifold is not a continuous attractor), but it is removed from the second largest (Fig. 4E & F).

所有解都涉及具有与任务中相关变量相同拓扑结构的慢流形. 不同的解在其渐近动力学上有所不同 (图 4). 最常见的解是不动点环流形类型 (图 4B 和 C). 这些解与观察到的持续活动依赖于离散吸引子一致. 较少见的拓扑包括围绕排斥环不变流形的慢环面 (图 4D). 这种解反过来与使用非恒定动力学进行记忆存储的可能性以及尽管存在时间变化活动但仍能维持持续表示的神经回路的观察一致. 所有稳定性结构 (不动点和极限环) 都映射到接近目标输出圆的位置 (图 S15、S19、S20). 我们还验证了图 4B 和 C 中训练网络的法向双曲. Jacobian 的最大特征值在零附近波动 (不变流形不是连续吸引子), 但它与第二大特征值相隔较远 (图 4E 和 F).

Universality amongst Good Solutions

The fixed point topologies show a lot of variation across networks (Fig. 4B,C, Fig. 5 and Fig. S20), much like the systems next to continuous attractors (Fig. 1 and Fig. 2). Previously, it has been observed that fixed point analysis has a major limitation, namely, that the number of fixed points must be equal across compared networks.

Our methodology effectively addresses and overcomes this limitation. The universal structure of continuous attractor approximations as slow invariant manifolds allows us to connect different topologies as approximate continuous attractors (Sec. 3.3). For results on LSTMs and GRUs and a higher dimensional task, see Sec. S7.7 and Sec. S7.9, respectively.

不动点拓扑在网络之间显示出很大的变化 (图 4B、C, 图 5 和图 S20), 就像连续吸引子旁边的系统一样 (图 1 和图 2). 此前已经观察到, 不动点分析有一个主要限制, 即 比较的网络之间的不动点数量必须相等.

我们的方法有效地解决并克服了这一限制. 作为慢不变流形的连续吸引子近似的普遍结构使我们能够将不同的拓扑结构连接为近似连续吸引子 (第 3.3 节). 有关 LSTM 和 GRU 以及更高维任务的结果, 请分别参见第 S7.7 节和第 S7.9 节.

Generalization Analysis

In this section, we use task-trained RNNs to study the relationship between dynamics and generalization capabilities. When neuroscientists study neural computations in animals, tasks have finite durations, leaving it unclear whether animals learn the intended computation or a finite-time approximation. The same issue applies to trained neural networks.

We will explore whether the networks possess the necessary memory for perfect recall or only perform the task within the timescale of their training.

在本节中, 我们使用任务训练的 RNN 来研究动力学与 泛化能力 之间的关系. 当神经科学家研究动物中的神经计算时, 任务具有有限的持续时间, 因此不清楚动物是否学习了预期的计算或有限时间的近似. 同样的问题也适用于训练过的神经网络.

我们将探讨这些网络是否具备完美回忆所需的记忆, 还是仅在其训练的时间尺度内执行任务.

The two possible approximations of a ring attractor, a limit cycle or a fixed point ring manifold (Sec. 3.3), exhibit markedly distinct generalization characteristics. Approximating the system as a limit cycle results in a memory trace that gradually diminishes over time (c.f. Park et al.).

Conversely, the alternative approximation's memory states are contingent upon the quantity and positioning of stable fixed points within the system. We describe in detail the generalization properties of the trained networks on the angular velocity integration task at two different time scales: asymptotic and finite time.

两个可能的环吸引子近似, 极限环或不动点环流形 (第 3.3 节), 表现出明显不同的泛化特性. 将系统近似为极限环会导致记忆痕迹随时间逐渐减弱 (参见 Park 等人).

相反, 另一种近似的记忆状态取决于系统中稳定不动点的数量和位置. 我们详细描述了训练网络在两个不同时间尺度下对角速度积分任务的泛化特性: 渐近时间和有限时间.

Figure 5: Temporal generalization validates theoretical predictions regardless of implementation detail. (A) Average accumulated angular error versus the maximum flow on the manifold (Eq. 5), shown for finite time (task duration that the networks were trained on, T1; filled markers) and at asymptotic time (hollow markers). (B) Normalized validation loss of all trained networks. (C) Average error and theoretical upper bound over time for two selected networks (corresponding to arrows in panel D). (D) Average asymptotic error is roughly inversely proportional to the number of fixed points. (E) Memory capacity is predictive of the average error.

图 5: 时间泛化验证了理论预测, 而不考虑实现细节. (A) 平均累积角误差与流形上的最大流动 (方程 5)之间的关系, 显示为有限时间 (网络训练的任务持续时间, T1; 填充标记)和渐近时间 (空心标记). (B) 所有训练网络的归一化验证损失. (C) 两个选定网络的平均误差和理论上限随时间变化 (对应于面板 D 中的箭头). (D) 平均渐近误差大致与不动点数量成反比. (E) 记忆容量可以预测平均误差.

Finite time

Along with the angular velocity integration component of the task, the trained networks learn to store a memory of an angular variable. We assess the performance of the network to store the memory of the angle over time. The networks typically perform well on the timescale on which they have been trained, T1 = 256 time steps (Fig. 5C). The memory error for T1 is, as theoretically predicted (Eq. 5, see Prop. 2), bounded by the uniform norm of the vector field on the invariant manifold, and therefore by the distance to a continuous attractor (Prop. 1, Fig. 5A, Sec. S7.6.3 & S6).

随着任务的角速度积分成分, 训练过的网络学会了存储角变量的记忆. 我们评估网络随时间存储角度记忆的性能. 网络通常在它们被训练的时间尺度 $T_{1} = 256$ 时间步上表现良好 (图 5C). $T_{1}$ 的记忆误差, 如理论预测 (方程 5, 见命题 2), 受不变流形上向量场的均匀范数限制, 因此受连续吸引子距离的限制 (命题 1, 图 5A, 第 S7.6.3 节和 S6 节).

Asymptotic time

Looking beyond the finite timescale provides valuable insights into the network's ability to store information. For the asymptotic time scale, we capture the asymptotic behavior of the system by identifying to what part of the state space the system evolves to in the limit $t\to\infty$ (see also Sec.S7.6.2). For a one-dimensional system, this will either be fixed points or a limit cycle. For the fixed-point type solution, the maximal error is given by the maximal distance to the next fixed point, while for a limit cycle, this will always be $\pi$. We calculate the average fixed point distance by taking the average of the inter-fixed-point interval for each neighboring pair of fixed points.

超越有限时间尺度提供了对网络存储信息能力的宝贵见解. 对于渐近时间尺度, 我们通过识别系统在极限 $t\to\infty$ 下演化到状态空间的哪一部分来捕捉系统的渐近行为 (另见第 S7.6.2 节). 对于一维系统, 这将是不动点或极限环. 对于不动点类型的解, 最大误差由到下一个不动点的最大距离给出, 而对于极限环, 这将始终为 $\pi$. 我们通过对每对相邻不动点的不动点间隔取平均值来计算 平均不动点距离.

We can also quantify the loss of information. Assuming a uniform distribution over the angles, we define the memory capacity as the negative conditional entropy of the continuous memory given the asymptotic state, i.e. the stable fixed points (see Sec. S7.6.2 and Eq. 68).

我们还可以量化信息的丢失. 假设角度上存在均匀分布, 我们将记忆容量定义为给定渐近状态 (即稳定不动点)的连续记忆的负条件熵 (见第 S7.6.2 节和方程 68).

Error Accumulation in Neural Networks

The mean accumulated error at the time at which the task was trained has an exponential relationship with the number of fixed points (Fig. 5A). Furthermore, this error is bounded by the mean distance between stable and unstable fixed points (red dots in Fig. 5D). This is another indication that the networks rely on a ring invariant manifold to implement the task. Networks with different numbers of fixed points might have the same performance on the finite time scale (bounded by $T_{1}||\varphi||_{\infty}$) but have vastly different generalization properties because they differ in the number of fixed points (Fig. 5C).

在任务训练的时间点, 平均累积误差与不动点的数量呈指数关系 (图 5A). 此外, 该误差受稳定和不稳定不动点之间的平均距离限制 (图 5D 中的红点). 这是另一个表明网络依赖于环不变流形来实现任务的迹象. 具有不同数量不动点的网络可能在有限时间尺度上具有相同的性能 (受 $T_{1}||\varphi||_{\infty}$ 限制), 但由于它们在不动点数量上存在差异, 因此具有截然不同的泛化特性 (图 5C).

Approximate Slow Manifolds are near Continuous Attractors

In Sec. 2.2, we presented a theory of approximate solutions in the neighborhood of continuous attractors. When are approximate solutions to the analog working memory problem near a continuous attractor? We posit that there are four conditions (see for more detail Sec. S5.3): (C1) sufficiently smooth approximate bijection between neural activity and memory content, (C2) the speed of drift of memory content is bounded, (C3) robustness against state (S-type) noise, and (C4) robustness against dynamical (D-type) noise 19. The correspondence implied by (C1) translates to the existence of a manifold in the neural activity space with the same topology as the memory content. Persistence (C2) requires that the flow on the manifold is slow and bounded. S-type robustness (C3) implies nonexpansive flow, i.e., non-positive Lyapunov exponents. Along with D-type robustness (C4), it implies the manifold is "attractive", and normally hyperbolic (see also Sec. S5.3.1).

在第 2.2 节中, 我们提出了连续吸引子邻域中近似解的理论. 什么时候模拟工作记忆问题的近似解接近连续吸引子? 我们假设有四个条件 (有关更多详细信息, 请参见第 S5.3 节): (C1) 神经活动与记忆内容之间的近似双射足够平滑, (C2) 记忆内容漂移的速度受限, (C3) 对状态 (S 型)噪声的鲁棒性, 以及 (C4) 对动力学 (D 型)噪声的鲁棒性. (C1) 所暗示的对应关系转化为神经活动空间中存在一个与记忆内容具有相同拓扑结构的流形. 持久性 (C2) 要求流形上的流动缓慢且受限. S 型鲁棒性 (C3) 意味着非扩张流动, 即非正 Lyapunov 指数. 连同 D 型鲁棒性 (C4), 它意味着流形是"有吸引力的", 并且是法向双曲的 (另见第 S5.3.1 节).

If these four conditions hold, for example for task-trained RNNs, there exists a smooth function with a uniform norm matching the slowness on the manifold such that when added, the slow manifold becomes a continuous attractor (Prop. 1 and Theorem 7, see also Sec. S5.4). For the RNN experiments, we added state-noise while training using stochastic gradient descent, satisfying (C3) and (C4). We have also verified that (C2) holds (Fig. 5A). Although the stochastic optimization cannot lead to the continuous attractor solution, it gets to the neighborhood where all approximate solutions share the same main feature: having a subsystem that has a slow flow.

如果这四个条件成立, 例如对于任务训练的 RNN, 存在一个平滑函数, 其均匀范数与流形上的缓慢性相匹配, 当添加时, 慢流形变为连续吸引子 (命题 1 和定理 7, 另见第 S5.4 节). 对于 RNN 实验, 我们在使用随机梯度下降训练时添加了状态噪声, 满足 (C3) 和 (C4). 我们还验证了 (C2) 成立 (图 5A). 尽管随机优化不能导致连续吸引子解, 但它进入了所有近似解共享相同主要特征的邻域: 具有一个具有慢流动的子系统.

Discussion

Continuous attractors are highly prone to bifurcation under arbitrary perturbations unless they exist in special parametric forms. This sensitivity to perturbations has traditionally made them seem unsuitable for modeling neural computation in noisy biological systems, according to conventional views on robustness. Nevertheless, we demonstrate that continuous attractors can exhibit functional robustness, making them a crucial concept in explaining the neural computation underlying analog memory. We show that approximations of analog memory (i.e., theoretical models that satisfy conditions (C1)-(C4)) must possess slow manifold dynamics, placing them near continuous attractors within the space of dynamical systems. This implies that both biological systems and artificial neural networks only need to be near a continuous attractor to effectively solve problems in a manner similar to the ideal theoretical model, on behaviorally relevant timescales.

连续吸引子在任意微扰下极易发生分岔, 除非它们以特殊的参数形式存在. 根据传统的稳健性观点, 这种对微扰的敏感性通常使它们看起来不适合在嘈杂的生物系统中建模神经计算. 然而, 我们证明了连续吸引子可以表现出功能稳健性, 使它们成为解释模拟记忆背后的神经计算的关键概念. 我们表明, 模拟记忆的近似 (即满足条件 (C1)-(C4) 的理论模型) 必须具有慢流形动力学, 将它们置于动力系统空间中接近连续吸引子的位置. 这意味着, 无论是生物系统还是人工神经网络, 只需接近连续吸引子, 就能在行为相关的时间尺度上以类似于理想理论模型的方式有效地解决问题.

Although we expressed our theory in a non-parametric manner with an arbitrary perturbation $\mathbf{p}(\cdot)$, we can easily extend it to particular parametric forms such as biophysical models or an RNN using a sensitivity of the flow to the parameters (e.g. synaptic weight). Our theory can be applied to latent dynamical systems estimated from neural recordings. As a framework, it can abstract out the details in imperfect dynamical implementations, however, it is an open problem to directly recover the continuous attractor from neural recordings or extend it to other ideal computational motifs.

尽管我们以非参数方式表达了我们的理论, 并使用任意微扰 $\mathbf{p}(\cdot)$, 但我们可以轻松地将其扩展到特定的参数形式, 例如生物物理模型或使用流动对参数 (例如突触权重) 的敏感性的 RNN. 我们的理论可以应用于从神经记录中估计的潜在动力系统. 作为一个框架, 它可以抽象出不完美动力学实现中的细节, 然而, 直接从神经记录中恢复连续吸引子或将其扩展到其他理想计算模式仍然是一个开放的问题.

Limitations

Although, we only explicitly describe the topology and dimensionality of the identified invariant manifolds for a representative set, the results indicate that most solutions have a ring invariant manifold with a slow flow. Our numerical analysis relies on identifying a time scale separation from simulated trajectories. If the separation of time scales is too small, it may inadvertently identify parts of the state space that are only forward invariant (i.e., transient). However, this did not pose a problem in our analysis of the trained RNNs, which is unsurprising, as the separation is guaranteed by state noise robustness (due to injected state noise during training).

尽管我们只明确描述了代表性集合中识别的不变流形的拓扑结构和维数, 但结果表明, 大多数解具有具有慢流动的环不变流形. 我们的数值分析依赖于从模拟轨迹中识别时间尺度分离. 如果时间尺度的分离太小, 它可能会无意中识别状态空间中仅向前不变 (即瞬态)的部分. 然而, 这在我们对训练 RNN 的分析中并没有造成问题, 这并不令人惊讶, 因为状态噪声鲁棒性 (由于训练期间注入的状态噪声)保证了分离.

The possible solutions that the networks can find are restricted by having a linear output mapping. In Park et al., an alternative dynamical solution using oscillators (or quasi-periodic toroidal attractor) was described, however, a nonlinear readout may be necessary.

网络可以找到的可能解受到线性输出映射的限制. 在 Park 等人中, 描述了使用振荡器 (或准周期环面吸引子)的替代动力学解, 但可能需要非线性读出.