3 Sequence Models

Classical ML/DL Practitioner Curriculum · 5:57

Listen on 93

Lyrics

[Verse 1]
RNN starts simple with a hidden state
Memory from yesterday, but gradients don't wait
They vanish as they travel through the chain
Long sequences forgotten, driving us insane
The basic cell just can't hold on too long
When backprop hits the layers, signals grow weak and wrong

[Chorus]
Remember the gates, forget and update
LSTM controls what flows and what waits
GRU simplified with reset and update gates
Attention mechanism says "look here, don't wait"
Convolution sliding through the temporal space
Three sequence models, each finding their place

[Verse 2]
LSTM came along with gates to save the day
Forget gate zeros out what should fade away
Input gate decides what new info to store
Output gate controls what flows to the fore
Cell state highway keeps the gradient strong
Now we can remember sequences long

[Chorus]
Remember the gates, forget and update
LSTM controls what flows and what waits
GRU simplified with reset and update gates
Attention mechanism says "look here, don't wait"
Convolution sliding through the temporal space
Three sequence models, each finding their place

[Verse 3]
GRU streamlined the gates from three down to two
Reset gate clears memory when starting anew
Update gate blends the past with present state
Fewer parameters but performance stays great
Bahdanau attention breaks the bottleneck curse
Looks at all encoder states, for better or worse

[Bridge]
But wait, there's more than recurrence alone
One-D convolution makes patterns known
Sliding filters capture local temporal features
Parallel processing, one of its best teachers
TCN stacks dilated convs in layers deep
Receptive fields growing, long patterns to keep

[Verse 4]
Luong attention came with different alignment
Global and local focus, perfect refinement
Dot product scoring or a learned function call
Weighted context vectors, attending to all
While recurrence struggles with parallel compute
Convolution processes sequences in one route

[Outro]
From vanishing gradients to attention's bright gleam
From LSTM gates to the convolutional dream
Each model has its strength in the sequence game
Understanding their purpose, that's how we claim
Mastery over time series, text, and speech
These three sequence models put learning within reach

← 3 Model Selection & Evaluation | 4 Practical Deep Learning →