AscendPrime Systems logoAscendPrime Systems

Production Engineering in Practice

Production Infrastructure Case Studies

Three production infrastructure challenges, followed by the architecture and operational practices used to engineer for them. The experience includes Redis Enterprise, ActiveMQ Artemis and MQTT messaging, WebRTC, and multi-region infrastructure.

Case studies

Production challenges, engineering decisions, and lessons learned.

Case Study A

Highly Available Redis Enterprise Infrastructure

Redis Enterprise infrastructure supporting approximately 20 applications with multi-zone high availability and operational rigor across 8 clusters and 90 databases.

Scale and environment

  • ~8 clusters
  • ~90 databases
  • ~20 applications
  • ~50,000 req/s
  • ~500 GB memory
Read engineering details

Environment and problem statement

The environment was production-facing and required a design that could sustain significant request volume while preserving failover readiness, operational clarity, and low-risk change management.

Challenge

The challenge was balancing throughput, failover behavior, deployment simplicity, and operational discipline across a large multi-database Redis footprint.

Key decisions and trade-offs

The design used rack and zone awareness to reduce the blast radius of infrastructure failures, while keeping data placement and failover dynamics aligned with production requirements. Monitoring, alerting, and runbooks were treated as first-class architectural controls so operators could respond quickly during degraded conditions.

Reliability, observability, and security

Availability and recovery decisions were shaped by failure-domain awareness, redundant capacity planning, and explicit operational procedures. Disaster recovery and recovery procedures were documented so recovery actions could be tested and repeated under pressure.

Testing and failure scenarios

The operating model included controlled failover exercises, monitoring validation, and recovery drills that checked whether the environment could tolerate broker loss and zone-level disruption without creating operational ambiguity.

Outcomes and lessons learned

The result was a Redis platform designed for resilient production operations at scale, with clear recovery paths and a strong emphasis on operational readiness rather than theoretical capacity alone.

Case Study B

MQTT Messaging Infrastructure at Scale

Apache ActiveMQ Artemis MQTT infrastructure designed for a large device footprint with layered security, connection control, and regional failover considerations.

Scale and environment

  • ~12 brokers
  • ~6 clusters
  • ~20,000 connected devices
  • 1 million device capacity target
  • mTLS + F5 + HA
Read engineering details

Environment and problem statement

A production messaging layer supporting a large number of connected devices needed to handle both connection density and operational resilience while preserving predictable behavior across distributed clusters.

Challenge

The design needed to accommodate device connection churn, broker distribution, and operational policies around TLS, authentication, and message ordering without creating hidden failure paths or uneven traffic distribution.

Key decisions and trade-offs

The architecture spread connectivity across broker clusters with F5-based load distribution, custom authentication handling, and client-IP logic to keep client connections balanced and predictable. Replication or mirroring, availability patterns, and failover routing were planned around the risk of connection storms and single-cluster dependency.

Reliability, observability, and security

Security and reliability were addressed together: mTLS protected device-to-broker trust, while connection distribution and failure handling minimized the impact of individual broker or cluster disruption. Operational risk reduction included deliberate review of client behavior, connection patterns, and recovery paths.

Testing and failure scenarios

Load testing focused on connection scale and distribution behavior, including how failures would surface under partial outage conditions. The work also reviewed ordering guarantees and operational quirks that appear only under sustained traffic and abnormal connection patterns.

Outcomes and lessons learned

The design produced an infrastructure model intended to scale toward a 1 million device target while keeping security, routing, availability, and operational recovery in active view.

Case Study C

Real-Time Communications Infrastructure

WebRTC, Coturn, and Janus infrastructure spanning multiple geographic regions to support camera deployments and production-grade connection establishment at scale.

Scale and environment

  • Multiple regions
  • WebRTC + Coturn + Janus
  • Camera deployments
  • NAT traversal involved
  • Regional distribution
Read engineering details

Environment and problem statement

The infrastructure needed to support real-time media traffic across multiple geographic areas while handling variable connectivity conditions and the realities of NAT traversal, signaling, and media path establishment.

Challenge

The main challenge was selecting connection patterns and regional placement that would keep sessions reliable while accounting for endpoint mobility, network traversal, and the availability demands of camera-based deployments.

Key decisions and trade-offs

The design centered on geographically distributed media infrastructure with explicit attention to how connection establishment, relay placement, and regional proximity affected availability and user experience. The architecture needed to be resilient enough to handle variable networking conditions without hiding operational complexity.

Reliability, observability, and security

Observability, region-aware placement, and failover planning were important because connection establishment and media path quality depend heavily on network conditions and the behavior of edge devices. Security and operational readiness were handled as part of the same infrastructure discipline.

Testing and failure scenarios

The work included validating the assumptions behind connection establishment, relay behavior, and region-based placement under stressed conditions rather than only ideal network conditions. This type of environment requires scenario testing around NAT traversal, regional degradation, and failure recovery.

Outcomes and lessons learned

The outcome was a realistic production-minded architecture pattern: distributed, capacity-aware, and operationally intentional rather than a single-region design optimized only for a happy path.

How we engineer

Architecture, observability, and resilience

Reference patterns and operational exercises for understanding how production systems behave under real-world conditions.

Architecture & reliability

Reference architectures

Illustrative reference architecture. Actual designs depend on workload, environment, operating model, and existing constraints.

Event-driven messaging flow

Illustrative reference architecture for producers, broker coordination, consumer groups, and observability.

ProducersLoadBalancerbootstrapBroker clusterTopics / partitionsReplicationISR / leadershipConsumer groupsMonitoring + alerting

High availability and disaster recovery

Multi-AZ deployment pattern with replication, monitoring, recovery orchestration, and a disaster-recovery region.

App tierZone APrimaryReplicationStandbyZone BMonitoring + recovery orchestrationDRRegionBackupRecovery

Secure messaging

Reference architecture for client identity, TLS/mTLS, authorization, broker infrastructure, and audit visibility.

ClientsmTLS / TLSIdentityAuthZ / ACLsBroker clusterRouting + replicationTopic-level policiesSecrets managementAudit loggingAudit trail

Observability, load testing & resilience

Production signals and failure scenarios

Illustrative demo data only. This is not connected to a live client environment.

Messaging operations dashboard

System health overview

Illustrative demo data
Consumer lag
8.4s
Group: payments-3
Msgs in / s
14.6k
Peak 22.4k
Broker CPU
72%
One broker trending high
URP / offline
2 / 0
Replica drift
Consumer lag by consumer group
Request latency + error rate
P95 latency174 ms
Error rate0.9%
Network throughput1.6 GB/s

Failure-injection exercises

  • Broker outage and leadership change

    Stop a message broker or node and observe how traffic and failover behavior react during a controlled degradation.

    • Failover events
    • Broker CPU and disk utilization
    • Replica lag
  • Consumer delay simulation

    Introduce slower processing and observe how backlog growth and recovery behave across a distributed workload.

    • Queue depth
    • Messages in/out per second
    • Rebalance events
  • Network interruption or AZ failure

    Simulate a network issue or availability-zone incident to observe failover behavior and request error rates.

    • Request latency
    • Error rate
    • Broker and network throughput
  • Client certificate revocation

    Revoke or expire a certificate to confirm that unauthorized clients are rejected before they can connect to message brokers.

    • Authentication failures
    • Connection counts
    • Audit log entries
  • Disaster recovery validation

    Exercise the documented recovery workflow to confirm that backup restoration, rehydration steps, and team procedures are still consistent.

    • Recovery time
    • Failover status
    • Operational checklist completion