Kafka Operations on Kubernetes
10 chapters
1. 5 Resource Tuning
[Verse 1]
When your Kafka cluster's running slow
CPU and memory need to flow
Set requests low but limits high
So brokers have room to reach the sky
ZooKeeper needs its resources too
Plan your limits, see performance through
[Chorus]
Tune it up, tune it down
Memory heap and threads all around
X-M-S for the start
X-M-X for the heart
Anti-affinity spread apart
Resource tuning is an art
[Verse 2]
JVM heap size sets the stage
Xms minimum for every age
Xmx maximum when load gets tough
G-one-G-C when standard's not enough
Garbage collection needs a plan
Parallel or CMS, choose your span
[Chorus]
Tune it up, tune it down
Memory heap and threads all around
X-M-S for the start
X-M-X for the heart
Anti-affinity spread apart
Resource tuning is an art
[Bridge]
Network threads handle the requests
I-O threads put storage to the test
Replica fetcher keeps data in sync
More threads help but use them smart, don't just think
[Verse 3]
Pod affinity brings pods close together
Anti-affinity spreads through any weather
Topology constraints across the zones
Keep your brokers on different nodes
Spread them wide for resilience sake
High availability's what we make
[Chorus]
Tune it up, tune it down
Memory heap and threads all around
X-M-S for the start
X-M-X for the heart
Anti-affinity spread apart
Resource tuning is an art
[Outro]
From CPU to GC tune
Network threads and memory soon
Spread your pods across the land
Strimzi tuning, take command
2. 3 Multi-Region / Disaster Recovery Patterns
[Verse 1]
When disasters strike your cluster down
Your data can't just disappear
Active-passive keeps a backup crown
One region leads, the other waits here
Primary serves all the traffic load
Secondary mirrors every byte
When the leader fails along the road
The passive wakes up for the fight
[Chorus]
Three patterns for recovery dreams
Active-passive, active-active schemes
RPO RTO guide your way
How much data loss, how long delay
Monitor that replication lag
Keep your Kafka systems flying the flag
Multi-region, disaster free
Strimzi patterns guarantee
[Verse 2]
Active-active runs them both alive
Two regions serving at the same time
But conflicts come when writes arrive
Need resolution paradigms
Last writer wins or timestamp rules
Vector clocks can solve the race
Custom logic with your tools
Keeps consistency in place
[Chorus]
Three patterns for recovery dreams
Active-passive, active-active schemes
RPO RTO guide your way
How much data loss, how long delay
Monitor that replication lag
Keep your Kafka systems flying the flag
Multi-region, disaster free
Strimzi patterns guarantee
[Bridge]
Recovery Point Objective shows
How much data you can lose
Recovery Time Objective knows
How long downtime you can choose
Milliseconds matter here
Watch the lag between your sites
JMX metrics crystal clear
Keep replication in your sights
[Verse 3]
Stretch clusters span across the zones
Single view but latency cost
Mirror Maker clones and owns
Cross-region data never lost
Choose your pattern by your need
Business rules will be your guide
Availability or speed
Let requirements help decide
[Chorus]
Three patterns for recovery dreams
Active-passive, active-active schemes
RPO RTO guide your way
How much data loss, how long delay
Monitor that replication lag
Keep your Kafka systems flying the flag
Multi-region, disaster free
Strimzi patterns guarantee
[Outro]
When the storms of failure rage
Your patterns keep you safe and sound
Turn another data page
Multi-region, battle-bound
3. Related KIPs
[Verse 1]
Five nine five starts the story
Raft protocol for metadata glory
Quorum leaders make decisions
No more Zookeeper revisions
Controllers talking through consensus
Building blocks that are tremendous
[Chorus]
KRaft evolution, six KIPs in motion
Five nine five, six thirty, seven sixty two
Eight three three, eight five three, nine sixty six too
Raft snapshots, rack aware election
Production ready, membership correction
Eligible leaders for the perfect connection
[Verse 2]
Six thirty brings us snapshots clean
Compact the log, keep memory lean
State gets saved in smaller chunks
No more growing storage junks
Bootstrap faster, recover quick
Snapshot magic does the trick
[Chorus]
KRaft evolution, six KIPs in motion
Five nine five, six thirty, seven sixty two
Eight three three, eight five three, nine sixty six too
Raft snapshots, rack aware election
Production ready, membership correction
Eligible leaders for the perfect connection
[Verse 3]
Seven six two knows your rack location
Controller spread across the nation
Fault tolerance gets much stronger
Downtime waits a little longer
Geography matters in the vote
Rack awareness keeps afloat
[Bridge]
Eight three three declares it's ready
Production workloads running steady
No more experimental warning
KRaft's bright new day is dawning
[Verse 4]
Eight five three changes membership
Add and remove without a skip
Dynamic clusters, scaling smooth
Controllers join and controllers move
Consensus adapts to cluster size
Growing and shrinking before your eyes
[Final Chorus]
KRaft evolution, six KIPs in motion
Five nine five, six thirty, seven sixty two
Eight three three, eight five three, nine sixty six too
Nine sixty six picks eligible leaders
Not every replica needs to be feeders
From Raft protocol to production readers
KRaft's the future for all believers
[Outro]
From metadata quorum to leader selection
KRaft's complete architectural perfection
4. 4 Logging
[Verse 1]
When your Kafka cluster starts to grow
Debug and info levels need to flow
Brokers humming with their chatter loud
Configure logging to cut through the crowd
Set the root logger to what you need
WARN for production, DEBUG to feed
Your troubleshooting when things go wrong
Structured data keeps the story strong
[Chorus]
Log it right, ship it tight
Fluentd flowing through the night
Elasticsearch storing all your traces
Loki keeping all the bases
Level up from INFO high
To ERROR when the systems cry
Structured logs in JSON form
Navigate through any storm
[Verse 2]
Connect workers need their own sweet tune
Set their levels or you'll debug soon
Operator logs will tell the tale
Of reconciliation without fail
Use log4j properties to control
What messages reach your monitoring goal
Filter noise but keep the signal clear
Performance metrics crystal clear
[Chorus]
Log it right, ship it tight
Fluentd flowing through the night
Elasticsearch storing all your traces
Loki keeping all the bases
Level up from INFO high
To ERROR when the systems cry
Structured logs in JSON form
Navigate through any storm
[Bridge]
Fluent Bit for lighter weight
When resources can't wait
Parse and forward every line
Timestamp correlation fine
Labels matter, tags are key
Search and query easily
Grafana dashboard shows
Where every error goes
[Verse 3]
Best practices keep you on the track
JSON structure, no looking back
Correlation IDs thread the story
From request start to final glory
Don't log secrets, keep them safe
Sampling helps when logs chafe
Retention policies save your space
While audit trails keep up the pace
[Final Chorus]
Log it right, ship it tight
Monitoring shining bright
Elasticsearch or Loki's grace
Both will give you searching space
DEBUG, INFO, WARN, ERROR too
Each level serves a purpose true
Structured logs in JSON form
Navigate through any storm
[Outro]
When Strimzi logs are flowing free
Observability is the key
From brokers to Connect and more
Logging opens every door
5. 1 Metrics with Prometheus
[Verse 1]
In your Kafka cluster running on K8s land
Strimzi gives you metrics right out of the box
JMX to Prometheus, conversion so grand
Built-in exporter that automatically talks
Configure metrics config in your Kafka resource
YAML declarations make the magic flow
No custom code needed, Strimzi's your source
For monitoring data you need to know
[Chorus]
Metrics, metrics, watch them flow
Under-replicated partitions, that's what you need to know
Latency, throughput, consumer lag too
Prometheus scraping, bringing insights to you
Monitor, measure, make it sing
Strimzi metrics, monitoring everything
[Verse 2]
Prometheus Operator deployed in your space
Custom resources managing your monitoring stack
ServiceMonitor pointing to the right place
Kafka metrics flowing, nothing held back
Configure the endpoint in metrics config section
Enable JMX exporter with a simple true
Target your Prometheus for metric collection
Real-time observability coming through
[Chorus]
Metrics, metrics, watch them flow
Under-replicated partitions, that's what you need to know
Latency, throughput, consumer lag too
Prometheus scraping, bringing insights to you
Monitor, measure, make it sing
Strimzi metrics, monitoring everything
[Bridge]
Request latency tells you response time story
Throughput shows messages per second rate
Under-replicated warns of partition worry
Consumer lag reveals if processing is late
Four key metrics keep your system healthy
Dashboard visualization makes problems clear
[Verse 3]
Grafana dashboards showing what's happening now
Alerts firing when thresholds are crossed
Strimzi makes monitoring simple somehow
No manual setup, no time lost
Built-in exporters handle the translation
From JMX beans to Prometheus format clean
Automatic service discovery configuration
Best observability you've ever seen
[Chorus]
Metrics, metrics, watch them flow
Under-replicated partitions, that's what you need to know
Latency, throughput, consumer lag too
Prometheus scraping, bringing insights to you
Monitor, measure, make it sing
Strimzi metrics, monitoring everything
[Outro]
From Kafka JMX to Prometheus store
Strimzi metrics give you so much more
Four key indicators, watch them all
Monitoring mastery, standing tall
6. 6 Lab: Cluster Upgrade & Rebalance
[Verse 1]
Starting with our cluster running strong and stable
Version N is working but we need to enable
The latest features that the new release will bring
Time to upgrade without breaking anything
Check the compatibility matrix first
Make sure dependencies won't burst
Rolling update is the safest way
Keep our services running every day
[Chorus]
Upgrade, rebalance, verify the flow
N to N plus one, watch the cluster grow
Add a broker, let Cruise Control decide
Partition distribution spread out wide
Upgrade, rebalance, verify the flow
Consumer groups stay stable as we go
[Verse 2]
Edit the cluster spec with the new version
Rolling restart begins, automated precision
Each pod comes down and up with latest code
Zero downtime following the upgrade road
Watch the operator logs for confirmation
Each broker joins without hesitation
Status shows us when the process completes
Successfully running on the latest release
[Chorus]
Upgrade, rebalance, verify the flow
N to N plus one, watch the cluster grow
Add a broker, let Cruise Control decide
Partition distribution spread out wide
Upgrade, rebalance, verify the flow
Consumer groups stay stable as we go
[Verse 3]
Now we scale up brokers, add one more node
Increase replicas in the cluster resource code
New broker starts but partitions stay the same
Cruise Control will optimize the data game
Create a rebalance proposal first
Analyze the plan before we burst
Check resource utilization goals
Balance network IO and storage loads
[Bridge]
Disk usage evenly spread
Network bandwidth properly fed
Replica count optimized
Leadership balanced and prized
[Verse 4]
Execute the rebalance plan with confidence
Watch partitions move with no turbulence
Consumer groups maintain their steady state
Lag metrics show they compensate
Verify the distribution looks just right
Each broker handling equal load tonight
Health checks passing, cluster running smooth
Version N plus one has found its groove
[Chorus]
Upgrade, rebalance, verify the flow
N to N plus one, watch the cluster grow
Add a broker, let Cruise Control decide
Partition distribution spread out wide
Upgrade, rebalance, verify the flow
Consumer groups stay stable as we go
[Outro]
From version N to N plus one we've climbed
Brokers balanced, perfectly timed
Strimzi cluster upgraded with care
Production ready, resilient and fair
7. 1 High Availability
[Verse 1]
When your Kafka needs to stay alive
Three brokers minimum to survive
One can fail but two remain
Keeping all your data streams in the game
Min in-sync replicas set to two
Guarantees your writes will make it through
No single point can bring you down
High availability wears the crown
[Chorus]
Three brokers strong, two in sync
PDBs so you never sink
Spread across the zones with care
Anti-affinity everywhere
High availability, that's the key
Kafka running endlessly
Three brokers strong, two in sync
That's the HA missing link
[Verse 2]
Pod disruption budgets guard your fleet
Kubernetes updates won't skip a beat
Maximum unavailable set to one
Rolling restarts won't leave you undone
Topology spread constraints take the stage
Distribute pods across every cage
Zones and regions, spread them wide
Failure domains can't coincide
[Chorus]
Three brokers strong, two in sync
PDBs so you never sink
Spread across the zones with care
Anti-affinity everywhere
High availability, that's the key
Kafka running endlessly
Three brokers strong, two in sync
That's the HA missing link
[Bridge]
Anti-affinity rules enforce the law
No two brokers on the same node raw
Required during scheduling time
Preferred when resources are prime
Topology keys define the scope
Hostname, zone, or region rope
Spread your brokers far and wide
Let availability be your guide
[Verse 3]
When disaster strikes a whole data center
Your design will be the dissenter
Cross-region spread keeps you alive
While others crash, you will thrive
Strimzi configs make it clean
Best availability you've ever seen
Three the magic number stays
For your always-running days
[Chorus]
Three brokers strong, two in sync
PDBs so you never sink
Spread across the zones with care
Anti-affinity everywhere
High availability, that's the key
Kafka running endlessly
Three brokers strong, two in sync
That's the HA missing link
[Outro]
Remember the rule of three and two
High availability will see you through
Spread and protect with budget care
Kafka's always running everywhere
8. 4 Lab: CDC Pipeline with Debezium
[Verse 1]
First we deploy our KafkaConnect cluster
Debezium PostgreSQL connector
Configure the database connection string
Watch the changes as they begin to sing
Source connector reads the transaction log
Every insert, update in the catalog
[Chorus]
CDC pipeline flowing free
Debezium captures what we need to see
From database to Kafka streams
Change data capture living the dream
Connect, capture, observe, and store
That's the pipeline we're building for
[Verse 2]
KafkaConnector resource we create
Point it to our sample database
Set the table whitelist configuration
Enable schema change documentation
Every row change becomes an event
To Kafka topics they are sent
[Chorus]
CDC pipeline flowing free
Debezium captures what we need to see
From database to Kafka streams
Change data capture living the dream
Connect, capture, observe, and store
That's the pipeline we're building for
[Bridge]
Watch the topics fill with data
Insert events and updates later
Delete operations marked as tombstones
Schema registry keeps us in the zone
Offsets tracked for exactly once
No duplicate events by happenstance
[Verse 3]
Now we add a sink connector too
S3 compatible storage for me and you
MinIO or AWS will do the trick
Parquet format makes queries quick
Events flow from source to destination
Complete pipeline configuration
[Chorus]
CDC pipeline flowing free
Debezium captures what we need to see
From database to Kafka streams
Change data capture living the dream
Connect, capture, observe, and store
That's the pipeline we're building for
[Outro]
From PostgreSQL to object store
Kafka Connect gives us so much more
Debezium makes the data flow
Now you've got the skills to go
9. 2 Performance Optimization
[Verse 1]
When your producers are moving too slow
Batch size matters, let the messages flow
Linger milliseconds, wait a little more
Pack them tight before they walk out the door
Compression's your friend when the data gets large
GZIP or LZ4, you're the one in charge
[Chorus]
Tune it up, tune it down
Performance secrets all around
Batch and linger, fetch and poll
Network threads take control
IOPS flying, GC's clean
Fastest Kafka you've ever seen
[Verse 2]
Consumer side needs its own special care
Fetch size bigger when you've got data to spare
Max poll records, don't bite more than you chew
Balance the load with what your app can do
Pull those messages in the perfect amount
Every millisecond's what we're here to count
[Chorus]
Tune it up, tune it down
Performance secrets all around
Batch and linger, fetch and poll
Network threads take control
IOPS flying, GC's clean
Fastest Kafka you've ever seen
[Bridge]
Broker level's where the magic happens
Network threads for the connections snapping
IO threads for the disk operations
Handle the load across all your nations
Eight network threads is usually right
Sixteen IO keeps your storage bright
[Verse 3]
Storage benchmarks tell the real story
IOPS and throughput, that's your glory
Sequential writes are what Kafka loves
Random reads when the consumer shoves
Test your disks before you deploy
SSD speed that you can enjoy
[Chorus]
Tune it up, tune it down
Performance secrets all around
Batch and linger, fetch and poll
Network threads take control
IOPS flying, GC's clean
Fastest Kafka you've ever seen
[Bridge 2]
G1 garbage collector's your best friend
Heap size matters from start to end
MaxGCPauseMillis set to twenty
Concurrent threads, you'll need plenty
JVM flags that keep it smooth
Performance gains that really soothe
[Outro]
From producer batch to consumer fetch
Every setting's got its perfect match
Network IO and storage speed
G1GC gives you all you need
Strimzi running like a dream
The fastest streaming you've ever seen
10. 6 Lab: Full Observability Stack
[Verse 1]
Start with Prometheus Operator installed and ready to deploy
Grafana waiting in the wings for dashboards we'll employ
The observability stack needs building from the ground
Full metrics and alerts where insights can be found
[Chorus]
Deploy, enable, import, trigger - that's the way we go
Prometheus grabs the metrics, Grafana makes them flow
Consumer lag alerts will tell us when things slow
Full observability stack - now watch your Kafka glow
[Verse 2]
Configure Strimzi cluster with metrics turned on bright
JMX exporter settings bring the data to light
Enable kafka metrics true and zookeeper the same
Consumer group lag tracking is the monitoring game
[Chorus]
Deploy, enable, import, trigger - that's the way we go
Prometheus grabs the metrics, Grafana makes them flow
Consumer lag alerts will tell us when things slow
Full observability stack - now watch your Kafka glow
[Bridge]
Import Strimzi dashboards into Grafana's embrace
Charts and graphs displaying cluster health and pace
Create consumer lag alert rules to watch the queue
Set thresholds that make sense for what you're trying to do
[Verse 3]
Now pause a consumer group to test the alert chain
Watch the lag build up as messages remain
The alert should fire when the threshold's been crossed
Observability saves the day when messages get lost
[Chorus]
Deploy, enable, import, trigger - that's the way we go
Prometheus grabs the metrics, Grafana makes them flow
Consumer lag alerts will tell us when things slow
Full observability stack - now watch your Kafka glow
[Outro]
From operator setup to alerts that really work
Full stack observability handles every quirk
Strimzi metrics flowing through the monitoring chain
Your Kafka cluster's health is crystal clear again
Back to Home