root_cause_category: replication_lag required_keywords: - replication lag - write-heavy workload - replica - wal - replay required_evidence_sources: - aws_cloudwatch_metrics - aws_performance_insights equivalent_root_cause_categories: - replication_lag_wal_volume forbidden_categories: - cpu_saturation - connection_exhaustion optimal_trajectory: - query_grafana_metrics - query_grafana_logs - query_grafana_alert_rules - describe_rds_instance - describe_rds_events max_investigation_loops: 4 model_response: | ROOT_CAUSE: The database hit replication lag because a write-heavy workload on the primary generated WAL faster than the read replica could replay it. The symptom is stale replica reads rather than a primary availability outage. ROOT_CAUSE_CATEGORY: replication_lag VALIDATED_CLAIMS: - ReplicaLag stayed above 900 seconds on payments-prod-replica-1. [evidence: aws_cloudwatch_metrics] - WriteIOPS and transaction log generation spiked on the primary during the same window. [evidence: aws_cloudwatch_metrics] - Performance Insights shows bulk UPDATE statements dominated DB load with WAL write waits. [evidence: aws_performance_insights] NON_VALIDATED_CLAIMS: - Reducing batch size or adding replica capacity would likely reduce the lag more quickly. CAUSAL_CHAIN: - A write-heavy workload drove sustained bulk updates on the primary. - WAL generation increased faster than the replica could replay. - Replication lag accumulated on payments-prod-replica-1 and stale reads triggered the alert.