Production Kafka
Learn how Kafka behaves in production — replication, ISR, KRaft, consumer lag, performance, monitoring, scaling, failure scenarios, and production architecture.
Format
Live, instructor-led remote sessions with hands-on labs
Price
$349
Next cohort
Oct 14, 2026 — October Cohort
Schedule
Thursdays, 10:00 AM-12:00 PM ET
Who it’s for
- Engineers operating Kafka clusters in production
- Platform and SRE teams responsible for messaging reliability
- Kafka Fundamentals graduates ready to go deeper
What you’ll leave able to do
- Explain how Kafka architecture decisions affect replication, leadership, and failure handling in production.
- Design for multi-AZ resilience and operational recovery without relying on a single happy path.
- Diagnose consumer lag, broker stress, and replication issues using monitoring and operational signals.
- Tune producer and consumer behavior for throughput and reliability under real workload conditions.
- Evaluate ISR, partition behavior, and message durability trade-offs in a production-minded design review.
- Use operational metrics and logs to identify failure modes before they become critical incidents.
- Apply disaster-recovery thinking to Kafka deployment patterns and broker loss scenarios.
- Assess the role of monitoring, observability, and recovery procedures in production Kafka operations.
Preview the learning experience
Sample lesson
Monitoring Consumer Lag and Degradation Signals
10-minute sample lesson
A short preview that demonstrates how lag, rebalance pressure, and throughput patterns are interpreted in practice.
- Identify lag by consumer group and partition.
- Review broker and topic health together.
- Correlate slow processing with operational metrics.
- Decide where to focus the next troubleshooting step.
Sample lab
Sample Lab: Kafka Failure and Recovery Review
Students simulate a broker loss, observe replication and leadership behavior, and review the recovery sequence used in production operations.
- Review the current cluster state and partition distribution.
- Stop a broker in a controlled lab environment.
- Observe leadership changes and replication status.
- Diagnose the impact on consumer lag and throughput.
- Document the recovery steps and expected operational outcome.
Curriculum
- 01Kafka architecture in production
- 02KRaft
- 03Brokers
- 04Partitions and replication
- 05ISR
- 06Failure scenarios
- 07Multi-AZ architecture
- 08Consumer lag
- 09Performance and throughput
- 10Capacity planning
- 11Monitoring and observability
- 12Kafka Connect
- 13Schema Registry
- 14Disaster recovery
- 15Production troubleshooting
- 16Architecture capstone
Hands-on labs
- Trigger and recover from a broker failure in a multi-broker cluster
- Tune producer and consumer configs for a throughput target
- Build a monitoring dashboard around the metrics that matter
- Design a disaster recovery plan for a sample architecture
- Capstone: design a production topology for a given workload
Instructor snapshot
Tulika Gupta
Tulika Gupta works with production infrastructure, distributed systems, and messaging concerns in real operational environments, with practical experience across Kafka, Redis, and high-availability systems.
Why this matters
- Production experience with Kafka and distributed messaging infrastructure in highly available environments.
- Operational understanding of failure domains, replication, observability, and incident response.
- Hands-on background across cloud operations, reliability engineering, and infrastructure architecture.
- Focus on practical decision-making under real production constraints rather than theory alone.
Format
Live, instructor-led remote sessions with hands-on labs
Prerequisites
- Kafka Fundamentals or equivalent experience performing practical Kafka work
- Comfort with distributed systems concepts and production operations
Enrollment, support, and policy
Payment & checkout
One-time payment for the cohort.
- Live instructor-led sessions
- Hands-on lab exercises and production-focused troubleshooting
- Course materials, architecture notes, and lab guidance
- Cohort-specific implementation discussions during the delivery period
Support
Community and post-course support options will be announced before the cohort begins.
- Live instruction during the cohort
- Production-focused labs and case discussions
- Instructor guidance within the delivery period
Enrollment policy
- Refund requests are considered on a case-by-case basis before the course start date and only when the enrollment terms have been explicitly approved by the founder.
- Students who withdraw before the course start date may be eligible for a credit or partial refund based on the approved policy and the amount of course preparation already incurred.
- If the instructor or cohort is rescheduled, students will be offered the next available cohort or a comparable alternative within the same course track when feasible.
- A cohort may be cancelled or rescheduled if minimum enrollment requirements are not met, and students will be notified in advance.
Draft policy — awaiting founder approval before publication
Upcoming Cohorts
October Cohort
Oct 14, 2026
January Cohort
Jan 20, 2027
Enrollment
Choose an open cohort
Selected cohort
October Cohort
- Production Kafka
- Thursdays, 10:00 AM-12:00 PM ET
- Oct 14, 2026
- Live Remote
- America/New_York
- Total: $349
Ready to start Production Kafka?