On September 14, 2026, between 15:30 and 17:10 UTC, Atlassian customers using Opsgenie, Jira Service Management, and Compass experienced delays in receiving alert notifications and intermittent failures when using Ops features. The issue was triggered by long-running transactions in a core service, which led to thread pool exhaustion and caused subsequent requests to become unresponsive. The long-running transactions were caused by a configuration change released between September 9, 2026 and September 11, 2026 UTC. The team actively monitored the impact of the change on the system until early September 14, 2026 UTC, and observed no anomalies. However, increasing traffic during US working hours on September 14, 2026, caused these unexpectedly long-running transactions. The incident was detected within one minute by our automated monitoring systems, and full service stabilization occurred after isolating and rolling back the change on September 14, 2026 at 17:10 UTC.
The incident primarily affected Ops features across Opsgenie, Jira Service Management, and Compass for customers hosted in the US region. During the incident, the end-to-end flow for alert creation and notification delivery experienced an average delay of 22 minutes. No alerts or notification payloads were dropped during the incident; all queued events were successfully processed and delivered as services recovered. Additionally, related Ops API endpoints and Web UI flows experienced elevated latency and intermittent errors. The disruption began at 15:30 UTC and was resolved by 17:10 UTC (total duration: 1 hour and 40 minutes).
Engineering teams were alerted immediately via backup disaster-alerting pipelines and began mitigation without delay. However, the disruption affected internal notification and coordination flows between cross-functional teams (including Customer Support and Incident Communications), which resulted in delays in publishing external Statuspage updates.
The incident was triggered when a newly released configuration change made redundant downstream service calls on every page load, causing a traffic spike. The configuration change, released between September 9, 2026, and September 11, 2026, was actively monitored for impact until early September 14, 2026. However, a traffic spike during US working hours caused unexpected issues. The downstream service began rate-limiting requests, and an aggressive retry strategy without sufficient backoff held worker threads open, leading to pool exhaustion and timeouts on incoming critical traffic.
The incident was mitigated by disabling the feature flag that triggered the inter-service rate limiting.
We understand reliable access to Atlassian products is critical for your teams. We are prioritizing the following actions to prevent recurrence:
We apologize to customers whose services were impacted during this incident; we are taking immediate steps to improve the platform’s performance and availability.
Thanks,
Atlassian Customer Support