Reliability issues with Jira and Confluence

Incident Report for Jira

Postmortem

Summary

On Jul 6, 2026, between 06:51 and 20:17 UTC, Atlassian customers using cloud products in European (EU) regions experienced service disruptions affecting automation rule execution, user search, user picker, and related workflows. The incident began when a core identity service experienced database saturation in EU regions, increasing latency and error rates. While responders worked to restore capacity, an emergency mitigation was applied which blocked automation rules from accessing the identity endpoint. In parallel, the elevated identity latency contributed to a cascading failure in a downstream user search service, degrading user search and user picker experiences across multiple products. Service was progressively restored after additional database read capacity was added, the emergency block was removed, and the user search service was manually scaled up.

Impact

The incident affected customers across multiple Atlassian products in EU regions on Jul 6, 2026 between 06:51 and 20:17 UTC.

  • Core Identity service degradation: elevated intermittent access denied rates were observed across products in EU regions.
  • Automation rule execution failures: automation rules using the default user (Automation for Jira) in Jira, Jira Service Management, and Jira Product Discovery in EU regions failed with a permissions error.
  • User search and user picker failures: user search success rates dropped significantly in EU regions, with user picker capability largely unavailable across products.

Root Cause

The incident originated from a recent configuration change that reduced the cache lifetime for a core identity service. As EU workday traffic ramped up on July 6, 2026, a larger share of requests began reaching the backing database directly rather than being served from cache. The EU databases did not have sufficient regional capacity to absorb the increased load, and database CPU reached saturation, causing elevated intermittent access denied rates.

As part of the mitigation, an emergency block was applied to reduce load on the saturated identity service database and prioritize restoring the core identity service. This inadvertently prevented automation rules from passing their pre-execution permission checks, causing them to fail. The extended duration was primarily due to a secondary wave of automation rule failures caused by the processing limit throttling.

In parallel, the sustained degradation caused increased latency and intermittent failures across products that depend on these checks, and contributed to a cascading failure in a downstream user search service. As a result, user search and user picker experiences became largely unavailable in EU regions until the permissions service recovered and the user search service was scaled up to restore capacity.

Remedial Actions Plan & Next Steps

We know outages impact customers' productivity. Atlassian is prioritizing the following actions to help prevent similar incidents in future:

  • Improve safeguards for high-impact emergency mitigations: Harden operational tooling so mitigation steps that could disable customer-facing workflows require additional review.
  • Strengthen identity infrastructure capacity and change safety: Review regional database sizing and capacity headroom to reduce the risk of saturation during peak traffic periods, and refine the assessment process for configuration changes so downstream load impact is evaluated before production rollout.
  • Strengthen resilience against cascading failures from upstream degradation: Audit backpressure handling, circuit breaker configuration, and autoscaling behavior in services, so that upstream degradation does not cause broader product impact across user search, user picker, and related experiences.
  • Harden replay mechanisms for Automation rule failures: Increase resilience of automation rules to reduce impact associated with incidents occurring in dependencies..

We recognize how critical reliable product workflows are for our customers, and we apologize to customers who were impacted by this incident.

Thanks,

Atlassian Customer Support

Posted Jul 17, 2026 - 16:59 UTC

Resolved

On July 6th 2026, between 06:51 UTC to 09:39 UTC some users may have experienced issues related Automation, and  between 08:10 UTC to 11:50 UTC some users may have experienced issues related to user-search, user-picker and assets in the EU-Region.

The issues have now been resolved and the services are operating normally for all the customers.
Posted Jul 06, 2026 - 13:04 UTC

Monitoring

We have applied the fix to mitigate the issue and are seeing recovery.

We continue to monitor the situation and should be fully recovered soon. We should be able to provide more updates in the next one hour or earlier.
Posted Jul 06, 2026 - 12:25 UTC

Investigating

We are still seeing some issues with user-search, user-picker and assets across with the EU-Region.

We continue to investigate the situation and will provide additional updates in the next one hour or sooner.
Posted Jul 06, 2026 - 11:46 UTC

Monitoring

We have successfully applied the fix and are seeing recovery. Automation capabilities are being restored and should operate normally going forward. Data processing requests submitted during the outage period were not completed due to service unavailability.

We continue to monitor the situation and expect full recovery within the next hour. We will provide additional updates in the next one hour or sooner.
Posted Jul 06, 2026 - 10:04 UTC

Investigating

We are investigating issues affecting your services.

We are actively working to identify the cause and resolve the issue. We will provide further updates within one hour.
Posted Jul 06, 2026 - 09:23 UTC
This incident affected: Viewing content, Create and edit, Authentication and User Management, Search, Notifications, Administration, Marketplace, Mobile, Purchasing & Licensing, Signup, and Automation for Jira.