Figure 1: Equivalent low-dimensional dynamics of diverse trained RNNs. (a) Schematic representation of the interval production task. The interval T determined by S1 and S2 is kept during a variable Tdelay, and reproduced upon a Go cue. Exemplary task-relevant population activity trajectories in state-space for a (b) N = 3 gated recurrent unit (GRU) network (including zoomed insets) and (c) PCA space for N = 25 unit GRU or N = 64 node vanilla RNN. In the trajectories, red lines denote S1 and S2 presentation, green - the Go cue, yellow - encoding and reproduction of T , purple - trajectory during Tdelay. Star/cross: start/end of the trajectory; circle: stable fixed point. Insets in (b,c): schematics of the respective RNN.
翻译:
图 1: 多样化训练的 RNN 的等效低维动力学。(a) 间隔产生任务的示意图。由 S1 和 S2 决定的间隔 T 在可变的 Tdelay 期间保持,并在 Go 提示下再现。状态空间中任务相关的群体活动轨迹示例,分别为 (b) N = 3 的门控循环单元 (GRU) 网络(包括放大插图)和 (c) N = 25 单元 GRU 或 N = 64 节点普通 RNN 的 PCA 空间。在轨迹中,红线表示 S1 和 S2 的呈现,绿色表示 Go 提示,黄色表示 T 的编码和再现,紫色表示 Tdelay 期间的轨迹。星号/叉号:轨迹的开始/结束;圆圈:稳定的固定点。(b,c) 中的插图:各自 RNN 的示意图。
Figure 2: Structured transient dynamics implement interval timing in trained RNNs. (a) Bifurcation diagrams of three exemplary trained N = 3 GRU networks. Blue / blacks lines: stable / unstable fixed points. SN: saddle-node bifurcation. Gray/red dashed lines: GRU solutions for input = 0 / input = 1. (b) Characterization of the state space during computations for the network depicted in (a, top) / Fig.1b, top row. Top row: The set of slow points/fixed points identified using the norm of the dynamics, q, for input = 0, including the flow visualized through randomly initialized trajectories (grey) is shown on the left, and the task relevant trajectory color-coded by its velocity in state space shown on the right. Circle / triangle: position of the fixed point / saddle-node bifurcation corresponding to (a). Bottom row: Maximal local eigenvalue (λ; zoom inset to the right) of the corresponding state-space structures. Remaining symbols as in Fig.1. For other examples see also Supp. Fig. S2. (c) Top: Snapshots of the temporal evolution of the network’s output for the example in (b) during S1 (left), in the T -encoding period (middle) and during Tdelay (right). Bottom: corresponding task-relevant trajectory segments and flow visualization. Remaining symbols and color coding as in (b). (d) Task-relevant trajectories for different T s for the example in (b), including zoomed representations (insets) of the trajectories’ segments during the delay period.
图 2: 结构化的瞬态动力学在训练的 RNN 中实现间隔计时。
(a) 三个示例训练的 N = 3 GRU 网络的分岔图。蓝色/黑色线:稳定/不稳定固定点。SN:鞍结分岔。灰色/红色虚线:输入 = 0 / 输入 = 1 的 GRU 解。
(b) 对 (a, top) / 图 1b, top row 中所示网络在计算期间的状态空间进行表征。 上排:使用动力学范数 q 确定的慢点/固定点集合,对于输入 = 0,包括通过随机初始化轨迹(灰色)可视化的流动显示在左侧,任务相关轨迹按其在状态空间中的速度进行颜色编码显示在右侧。圆圈/三角形:对应于 (a) 的固定点/鞍结分岔的位置。
下排:相应状态空间结构的最大局部特征值 (λ;右侧放大插图)。其余符号如图 1 所示。其他示例请参见补充图 S2。
(c)
上排:在 (b) 中的示例中,网络输出在 S1(左)、T 编码期间(中)和 Tdelay(右)期间的时间演变快照。
下排:相应的任务相关轨迹段和流动可视化。其余符号和颜色编码如 (b) 所示。
(d) 在 (b) 中的示例中,不同 T 的任务相关轨迹,包括延迟期间轨迹段的放大表示(插图)。
Figure 3: Emergence of sets of slow points during GRU training. (a) Representative loss curve of the initial 50 trained epochs with a zoomed inset of the full curve (left) and the evolution of the bifurcation structure of the network during training for the marked epochs (0, 6 and 21, right). Blue / black lines: stable / unstable fixed points; gray dashed line: organization at input = 0. The example corresponds to Fig. 2a, top. (b) Snapshots of the task-relevant trajectories and the sets of slow points (top), as well as the corresponding local maximal eigenvalue (bottom) at the marked epochs during training. Trajectory color-coding as in Fig. 1b. For additional example, see Supp. Fig. S4a,b. (c) Bifurcation diagram of a trained 3-unit GRU network with respect to a bias parameter bh12 (line color as in (a)) and the respective timing error calculated for different bh12 values (red line). Vertical black/grey dashed line: position of the SN bifurcation / organization after training. Related to Supp. Fig. S4c.
图 3: GRU 训练期间慢点集合的出现。
(a) 代表性的初始 50 个训练时期的损失曲线,带有完整曲线的放大插图(左)以及在标记时期(0、6 和 21)期间网络分岔结构的演变(右)。蓝色/黑色线:稳定/不稳定固定点;灰色虚线:输入 = 0 时的组织。该示例对应于图 2a, top。
(b) 在训练期间标记时期的任务相关轨迹和慢点集合的快照(上),以及相应的局部最大特征值(下)。轨迹颜色编码如图 1b 所示。有关其他示例,请参见补充图 S4a,b。
(c) 针对偏置参数 bh12 的训练 3 单元 GRU 网络的分岔图(线颜色如 (a) 所示)以及针对不同 bh12 值计算的相应计时误差(红线)。垂直黑色/灰色虚线:SN 分岔的位置/训练后的组织。与补充图 S4c 相关。
Figure 4: Generalization property of trained networks relies on the extent of the set of slow points. (a) Response timeseries of a N = 3 unit GRU network trained to encode T = 80a.u., and tested on T = 60a.u. when the Go cue amplitude is equal to 1 (top), and 0.172 (bottom - corresponding to blue cross in (c)). Solid black/cyan (red) lines: input/output timeseries, ∆T - timing error in response. (b) The set of slow points of the trained network characterized by q and the population activity trajectories for the two Go cue amplitudes shown in (a). Trajectories color coded as in Figure 1b. Dashed/solid lines correspond to Go cue amplitude 1/0.172 respectively. Related to Supp. Fig. S6a,b. (c) Identified Go cue amplitudes and corresponding ∆T quantification resulting in generalization for encoding 20 < T < 100. (d) Generalization extent of the network in Fig.2a - top. Solid black/red line: optimal / actual predicted T . Blue line: Time spent by the trajectory during the memory phase (delay period). Gray shaded area: T interval presented during network training. Black/red dashed lines correspond to T = 115/140a.u. respectively, shown in (e). Additional example shown in Supp. Fig. S6c,d. (e) Task-dependent trajectories and set of slow points for T = 115a.u. (left) and T = 140a.u. (right).
图 4: 训练网络的泛化特性依赖于慢点集合的范围。
(a) 针对 T = 80a.u. 进行编码训练的 N = 3 单元 GRU 网络的响应时间序列,并在 Go 提示幅度等于 1(上)和 0.172(下 - 对应于 (c) 中的蓝色交叉点)时测试 T = 60a.u.。实线黑色/青色(红色)线:输入/输出时间序列,响应中的 ∆T - 计时误差。
(b) 训练网络的慢点集合由 q 表征,以及 (a) 中显示的两个 Go 提示幅度的群体活动轨迹。轨迹颜色编码如图 1b 所示。虚线/实线分别对应于 Go 提示幅度 1/0.172。与补充图 S6a,b 相关。
(c) 确定的 Go 提示幅度和相应的 ∆T 量化,导致编码 20 < T < 100 的泛化。
(d) 图 2a - top 中网络的泛化范围。实线黑色/红色线:最佳/实际预测 T。蓝线:轨迹在记忆阶段(延迟期)所花费的时间。灰色阴影区域:网络训练期间呈现的 T 间隔。黑色/红色虚线分别对应于 T = 115/140a.u.,如 (e) 所示。补充图 S6c,d 显示了其他示例。
(e) T = 115a.u.(左)和 T = 140a.u.(右)的任务相关轨迹和慢点集合。
Figure 5: Dynamical systems model for flexible temporal computations. (a) Schematic representation of the model: G1 − G3 - single ghost variables, x - timing variable, z - memory variable (details in Methods). Arrows/block arrows: activation/inhibition. (b) Exemplary time series depicting the model response for T = 250a.u. Black dashed line: dx/dt as a function of time. (c) Estimated T interval range for which the model achieves < 15% error in prediction.
图 5: 灵活时间计算的动力系统模型。
(a) 模型的示意图:G1 − G3 - 单个幽灵变量,x - 计时变量,z - 记忆变量(方法中有详细说明)。箭头/块箭头:激活/抑制。
(b) 描述模型响应的示例时间序列,针对 T = 250a.u.。黑色虚线:dx/dt 随时间的函数。
(c) 估计模型在预测中实现 < 15% 误差的 T 间隔范围。