[Verse 1] RNN starts simple with a hidden state Memory from yesterday, but gradients don't wait They vanish as they travel through the chain Long sequences forgotten, driving us insane The basic cell just can't hold on too long When backprop hits the layers, signals grow weak and wrong [Chorus] Remember the gates, forget and update LSTM controls what flows and what waits GRU simplified with reset and update gates Attention mechanism says "look here, don't wait" Convolution sliding through the temporal space Three sequence models, each finding their place [Verse 2] LSTM came along with gates to save the day Forget gate zeros out what should fade away Input gate decides what new info to store Output gate controls what flows to the fore Cell state highway keeps the gradient strong Now we can remember sequences long [Chorus] Remember the gates, forget and update LSTM controls what flows and what waits GRU simplified with reset and update gates Attention mechanism says "look here, don't wait" Convolution sliding through the temporal space Three sequence models, each finding their place [Verse 3] GRU streamlined the gates from three down to two Reset gate clears memory when starting anew Update gate blends the past with present state Fewer parameters but performance stays great Bahdanau attention breaks the bottleneck curse Looks at all encoder states, for better or worse [Bridge] But wait, there's more than recurrence alone One-D convolution makes patterns known Sliding filters capture local temporal features Parallel processing, one of its best teachers TCN stacks dilated convs in layers deep Receptive fields growing, long patterns to keep [Verse 4] Luong attention came with different alignment Global and local focus, perfect refinement Dot product scoring or a learned function call Weighted context vectors, attending to all While recurrence struggles with parallel compute Convolution processes sequences in one route [Outro] From vanishing gradients to attention's bright gleam From LSTM gates to the convolutional dream Each model has its strength in the sequence game Understanding their purpose, that's how we claim Mastery over time series, text, and speech These three sequence models put learning within reach
← 3 Model Selection & Evaluation | 4 Practical Deep Learning →