Status updates | Cerebrium Incidents and maintenance reported on status page for Cerebrium https://status.cerebrium.ai/ https://d1lppblt9t2x15.cloudfront.net/logos/31d4d8d67ce2324501b36dac1cf60977.png Status updates | Cerebrium https://status.cerebrium.ai/ en Slow container image pulls in eu-north1 https://status.cerebrium.ai/incident/1056990 Tue, 08 Sep 2026 16:42:00 -0000 https://status.cerebrium.ai/incident/1056990#f126c84761fd0771a515d525d8b96d1ffbcd97d16c27ccc46e69210cdbe2ccf2 Incident The issue has been resolved. Slow container image pulls in eu-north1 https://status.cerebrium.ai/incident/1056990 Tue, 08 Sep 2026 15:36:00 -0000 https://status.cerebrium.ai/incident/1056990#805a948e3ac6b96d97145ebd306bc0676e84da3af23a6b8f4f49acb7b6511e37 Incident We're seeing elevated S3 latency in eu-north1, which is slowing container image pulls. New deploys and cold starts in that region are taking longer than usual. Inference is unaffected: apps with warm replicas are serving traffic, and other regions are running normally. We're currently investigating and will post an update shortly. Maintenance: Scheduled provider maintenance: delayed GPU node provisioning in US East (Crusoe) and US South Central (Crusoe https://status.cerebrium.ai/maintenance/1056631 Tue, 08 Sep 2026 09:08:10 -0000 https://status.cerebrium.ai/incident/1056631#55bef11f62c8d90a5847e2f61dcb1740d8f002eec6eb3e5af4c61fafd7bf1a4f Maintenance Our infrastructure provider Crusoe is performing maintenance on the Kubernetes control plane for the clusters behind our US East and US South Central Crusoe regions. What to expect: - Running apps and live traffic are not affected. - Scaling onto GPUs that are already provisioned continues to work normally. - Requests to provision new GPU nodes may fail for 5 to 10 seconds at a time and then retry. If your app needs to scale beyond the GPUs currently available in the region during this window, new replicas may take a few minutes longer than usual to become ready and requests may queue briefly. - Apps with cross-region enabled will spill over automatically as normal. What you can do: - Avoid deploying new versions of your apps in this window if you can. A deploy runs old and new replicas side by side and is the most common reason an app needs extra nodes. - If you have a known traffic peak in this window, consider raising your minimum replicas or scaling buffer from about 16:15 UTC so capacity is warm before the maintenance starts, and lowering it again afterwards. Provider network outage: Nebius eu-north1 https://status.cerebrium.ai/incident/1046716 Wed, 02 Sep 2026 17:17:00 -0000 https://status.cerebrium.ai/incident/1046716#20228521a02eb47d1dc5270b831506fd39f6c9fce9442253e649ef03dc117a08 Incident The issue has been resolved and the provider and region are stable again. Provider network outage: Nebius eu-north1 https://status.cerebrium.ai/incident/1046716 Wed, 02 Sep 2026 09:57:00 -0000 https://status.cerebrium.ai/incident/1046716#8b06ddaffa5c6fc5010aee9660c4c12989a04ad80b2608724999e254a27e6824 Incident We are currently tracking intermittent network outages with an upstream provider (Nebius) in the eu-north1 region. We are working with provider to identify and resolve the issue and will provide updates as the provider investigations progress. Upstream outage: us-central1 https://status.cerebrium.ai/incident/1021177 Wed, 19 Aug 2026 16:26:00 -0000 https://status.cerebrium.ai/incident/1021177#4a93dc5c0e07c66275636a5ecbe7e68b48c47e5ffef59c58173f04386fe1ef70 Incident Upstream provider has resolved the issue. Upstream outage: us-central1 https://status.cerebrium.ai/incident/1021177 Wed, 19 Aug 2026 12:59:00 -0000 https://status.cerebrium.ai/incident/1021177#84014473e1644e1d7ad594d173174b821d94a2a9de16104d0c7c703541d909a7 Incident The issue has been resolved upstream and services should resume shortly Upstream outage: us-central1 https://status.cerebrium.ai/incident/1021177 Wed, 19 Aug 2026 11:37:00 -0000 https://status.cerebrium.ai/incident/1021177#498e93ebc1630e33c95d0c636eb62be3539243979b4f854989f5a59246ee714b Incident The root cause has been identified as hardware related degradation. GPU performance in the region is impacted and inference to containers in the Nebius us-central1 region currently suffers from intermittent failure. Upstream outage: us-central1 https://status.cerebrium.ai/incident/1021177 Wed, 19 Aug 2026 11:19:00 -0000 https://status.cerebrium.ai/incident/1021177#8dcf3644130b205d4bd997d23add06e13fc51710be5365a525ec225cc37b04c3 Incident We are currently investigating an upstream issue with our Nebius provider in us-central1. We will provide regular updates as more information becomes available Maintenance: Emergency maintenance: Inference gateway https://status.cerebrium.ai/maintenance/1009423 Wed, 12 Aug 2026 09:55:17 -0000 https://status.cerebrium.ai/incident/1009423#e943943a0d300e231ca710d7c5a10cb3ab21e9b193296a194255dd99da933c3e Maintenance We are performing emergency maintenance on our gateway infrastructure to improve reliability and capacity headroom. During the window, active connections will be dropped and in-flight requests may be terminated. New requests may fail for up to 30 seconds while the gateway cycles. Clients with standard retry logic should recover automatically. What to expect: - Brief connection drops during the window - Possible request failures for up to 30 seconds - No action needed if your integration retries failed requests What we recommend: - Avoid starting long-running requests just before or during the window - Confirm your client retries on connection errors - We'll post updates here as the maintenance progresses and confirm once complete. Networking issues https://status.cerebrium.ai/incident/994402 Mon, 03 Aug 2026 21:28:00 -0000 https://status.cerebrium.ai/incident/994402#e404d6dc34103d3bbf034459fb0ea588bd55b8ebdefdc72f654a75892b73a04b Incident Up stream provider issues have been resolved and traffic should Networking issues https://status.cerebrium.ai/incident/994402 Mon, 03 Aug 2026 19:33:00 -0000 https://status.cerebrium.ai/incident/994402#223a1b822ffa09da607852469deab0fd04109b222a4f529d004e7b765cb4e2df Incident The issue has been mitigated for now. We are continuing to monitor and are preparing to reinstate the cluster. Networking issues https://status.cerebrium.ai/incident/994402 Mon, 03 Aug 2026 18:30:00 -0000 https://status.cerebrium.ai/incident/994402#b1336fc41620f739e065b19613a1ed8190c60552e5fde0e3e60317ff5dd0fe0f Incident We're currently experiencing networking issues in US East affecting one of our us-east clusters. Maintenance: Scheduled Maintenance: Global router upgrade and Nebius image rollouts https://status.cerebrium.ai/maintenance/988738 Fri, 31 Jul 2026 13:29:34 -0000 https://status.cerebrium.ai/incident/988738#bdf8a63ae0efc05c12809362be796eed34e032d1def86c55e50367bce5819425 Maintenance We are upgrading our global router to improve latency and routing decisions, and rolling out new machine images across our Nebius regions to enable checkpointing. Requests routed through the global router may see increased latency or transient errors while the upgrade rolls out. Nodes in Nebius eu-north1 and us-central1 will be rolled onto the new image. Active runs in these regions may be interrupted. Regions outside Nebius are not affected by the image rollout. Deployments and inference requests may still see elevated latency across all regions while the router work completes. If you run latency-sensitive production traffic, plan for retries during the window. We will post updates here as the work progresses. Registry unavailable in EU-North1 https://status.cerebrium.ai/incident/979395 Sun, 26 Jul 2026 17:32:00 -0000 https://status.cerebrium.ai/incident/979395#05437c25cdce3542d5a45557482ae2b57ce7ae5a9bcd70a73afa23c3c67dadc5 Incident Normal operation has resumed. Registry unavailable in EU-North1 https://status.cerebrium.ai/incident/979395 Sun, 26 Jul 2026 17:27:00 -0000 https://status.cerebrium.ai/incident/979395#61f7f5d7946cf9d452e6e2093dfd3a9d087163af8ff6179e854e1a899517313d Incident We are receiving and investigating reports of issues with our Registry in the Nebius EU-North1 cluster. Maintenance: Planned maintenance in Crusoe provider https://status.cerebrium.ai/maintenance/973297 Wed, 22 Jul 2026 08:18:02 -0000 https://status.cerebrium.ai/incident/973297#9e4659fc3f9959c7213e179161a8b51a551dc1d64c685f6cc8bb1d31586d64ee Maintenance One of our upstream infrastructure providers will be conducting planned network maintenance in the us-southcentral region. No customer action is required. Workloads will be moved to alternative capacity automatically ahead of the window. You should see no change in availability or performance. All other regions remain fully operational. If you notice any issues during or after the window, please reach out to us. Inference downtime https://status.cerebrium.ai/incident/971924 Tue, 21 Jul 2026 17:55:00 -0000 https://status.cerebrium.ai/incident/971924#8fa646eaef041e9e3d432392ea49412537ba8526916b35bdbe903c078cfe1c34 Incident The issue has been resolved and all regions are fully operational again. Inference downtime https://status.cerebrium.ai/incident/971924 Tue, 21 Jul 2026 17:45:00 -0000 https://status.cerebrium.ai/incident/971924#fa2183ebcb35e98ff3664644a1f48ece41140b5b6dbd768a97dad2851d266f1b Incident Inference requests to all regions are 4xx'ing New builds & inference in US are currently affected by ongoing maintenance downstream https://status.cerebrium.ai/incident/966057 Thu, 16 Jul 2026 20:37:00 -0000 https://status.cerebrium.ai/incident/966057#1e072545b8635e240bc66818d00d1b0a343edaabedd6cf9398155c80ac1932c4 Incident Maintenance has been completed and the region is fully restored. Expect a few slower cold starts while caches warm up. New builds & inference in US are currently affected by ongoing maintenance downstream https://status.cerebrium.ai/incident/966057 Thu, 16 Jul 2026 17:24:00 -0000 https://status.cerebrium.ai/incident/966057#a20e95681fe21d57eacdd0f0f90e53756c5c91243a22e4439a628f6b4b760c6d Incident Maintenance in the us-east1-a region is complete. Our upstream provider is still restoring access to a small number of remaining hypervisors, expected to be fully recovered by 17:45 UTC, at which point (and after verification) access to the region will be restored. Customers on our global router remain unaffected (Insofar as inference is concerned). Traffic has been failing over to other regions and providers throughout, no action required. Builds continue to remain affected If you have workloads pinned directly to us-east1-a that are still unavailable, they should recover automatically. You can also pin to a different provider or region in the meantime. New builds & inference in US are currently affected by ongoing maintenance downstream https://status.cerebrium.ai/incident/966057 Thu, 16 Jul 2026 12:38:00 -0000 https://status.cerebrium.ai/incident/966057#6787b107ade52f8f41efdeb55a400d27d1a0d7b28841b6dd70bac0fcfae3f6f6 Incident Planned maintenance is affecting new builds in the US from starting. We expect this to be resolved shortly. Inference and other services are not affected. US Partial Outage - Mostly effecting new deployments of L40s https://status.cerebrium.ai/incident/965239 Thu, 16 Jul 2026 02:10:00 -0000 https://status.cerebrium.ai/incident/965239#4573e969461dc2741150ca8ac8235dc625ac36318daa3a251bbed7161c83e58a Incident Services are operating normally again US Partial Outage - Mostly effecting new deployments of L40s https://status.cerebrium.ai/incident/965239 Thu, 16 Jul 2026 00:06:00 -0000 https://status.cerebrium.ai/incident/965239#2ba9123b9a113ee1103db81716f9bad83e6c48d9eaaf6e9b6989e3d6d8c356a9 Incident Some issues persist with that provider US Partial Outage - Mostly effecting new deployments of L40s https://status.cerebrium.ai/incident/965239 Wed, 15 Jul 2026 23:55:00 -0000 https://status.cerebrium.ai/incident/965239#ff42b908754164a500627b0227de932b753f19e6244502a8884321013f94ad84 Incident Underlying provider has resolved the issue US Partial Outage - Mostly effecting new deployments of L40s https://status.cerebrium.ai/incident/965239 Wed, 15 Jul 2026 23:40:00 -0000 https://status.cerebrium.ai/incident/965239#378427a9a13772ba53bda9f9a19bdf53af2860307fb12344cf5e388893a0d09f Incident 1 provider is down in the US East region. Existing apps should shift to available providers and regions. New deployments of L40s will be effected until the issue is resolved. Maintenance: Planned maintenance - Intermittent availability in us-east-1 https://status.cerebrium.ai/maintenance/963404 Tue, 14 Jul 2026 07:37:22 -0000 https://status.cerebrium.ai/incident/963404#06df9b086028fa8f16a9a12d7829ab64b0874e6c3dae53121030cf7768170574 Maintenance One of our upstream infrastructure providers (Crusoe) will be conducting planned network maintenance affecting the us-east-1 region. Most customers will not be affected. If you're on our global router, your workloads will automatically fail over to a different region and/or provider; no action is required. If you're running workloads pinned directly to Crusoe us-east-1, you may experience intermittent unavailability during the maintenance window. To avoid disruption, you can pin your application to a different provider or region for the duration, or take advantage of our global routing to handle failover automatically. All services in other regions will remain fully operational. US East 1 partial outage - Traffic shifted https://status.cerebrium.ai/incident/948344 Thu, 09 Jul 2026 15:59:00 -0000 https://status.cerebrium.ai/incident/948344#25feed2b4c65e9214e287048b516ee3dc9cc668c3c62311a2bc6c4654a122267 Incident The issue has been patched upstream and Cerebrium health checks have been succeeding for the past 45 minutes. Marking as resolved. US East 1 partial outage - Traffic shifted https://status.cerebrium.ai/incident/948344 Thu, 09 Jul 2026 15:06:00 -0000 https://status.cerebrium.ai/incident/948344#f413c0243517409117b53b6825f1f088188995e0c169f7b39561c513fdd2b654 Incident The issue is on going and we are still communicating with the upstream provider. US East 1 partial outage - Traffic shifted https://status.cerebrium.ai/incident/948344 Thu, 09 Jul 2026 14:55:00 -0000 https://status.cerebrium.ai/incident/948344#9e2833e16630fe2b03d1546c2866beed312f635ebb543452ffbb82e8827c9441 Incident The upstream provider has implemented a mitigation. Services should be restored, however we are still tracking this issue and will revert with any updates. US East 1 partial outage - Traffic shifted https://status.cerebrium.ai/incident/948344 Thu, 09 Jul 2026 14:26:00 -0000 https://status.cerebrium.ai/incident/948344#ea78e5fcdd3e396a6b6f2ea2c2819f2bbc27ace253027ea41e1e16ab8e362188 Incident 1 provider in US East 1 is having networking issues. Traffic should automatically reroute to others but service is degraded New Builds Not Starting https://status.cerebrium.ai/incident/944943 Mon, 06 Jul 2026 14:52:00 -0000 https://status.cerebrium.ai/incident/944943#91ddf7738d0855d164b04f6f46ffb0e2f7bf401f8dcfbaade5a7c54e0f0f4fb8 Incident The issue has been resolved and builds are back up and running again. New Builds Not Starting https://status.cerebrium.ai/incident/944943 Mon, 06 Jul 2026 14:42:00 -0000 https://status.cerebrium.ai/incident/944943#5ab281ed0a1f0f93f38ecf5d2e7c49a8758671e88f9f415680fed97addea9477 Incident New builds are currently stalling. Existing apps are running fine. Degraded registry performance in AWS us-east-1 https://status.cerebrium.ai/incident/940412 Wed, 01 Jul 2026 13:39:00 -0000 https://status.cerebrium.ai/incident/940412#3d5c1d369188df1ff597d8882e5fa7bd46b8fc3bb0e7334a4aa21609c3f8e655 Incident We've identified the root cause of the issue as constrained registry resources. These resources have been scaled and the issue has been resolved Degraded registry performance in AWS us-east-1 https://status.cerebrium.ai/incident/940412 Tue, 30 Jun 2026 20:00:00 -0000 https://status.cerebrium.ai/incident/940412#3f8e78c72ea4217519df568c568351f6ed096c8cef183aabdfc91ec0c6d95fd2 Incident We're currently experiencing degraded registry performance do to an issue with a service related to an upstream provider. We are currently investigating and will provide an update as soon as we have been able to identify the cause. Maintenance: Scheduled Maintenance: Crusoe us-east-1 region unavailable https://status.cerebrium.ai/maintenance/924585 Mon, 15 Jun 2026 15:26:10 -0000 https://status.cerebrium.ai/incident/924585#8446f1858f340618ffa9470163c92b674d7f208934e3b574354799ca54c66165 Maintenance We are performing urgent maintenance in our Crusoe us-east-1 region tomorrow, from 10:00 - 16:00 UTC. Apps pinned to this region and provider will be affected for the duration of the window. To avoid interruption, please deploy to another Cerebrium region or provider before 10:00 UTC tomorrow. Apps running on our new global infrastructure are unaffected. They will fail over automatically to other regions/providers - no action is required. Thank you for your continued support and patience. Builds are broken https://status.cerebrium.ai/incident/917975 Mon, 08 Jun 2026 13:49:00 -0000 https://status.cerebrium.ai/incident/917975#0d37513e15e40b0cecbb0c3a7d7c562a980c49fbbd50b35ae6db78639df3360e Incident Build service is fully restored Builds are broken https://status.cerebrium.ai/incident/917975 Mon, 08 Jun 2026 13:43:00 -0000 https://status.cerebrium.ai/incident/917975#c3e032c254535c77f650c49aa5d299f35f671109a445fc87c80dcbcc037d3b7e Incident Most builds are working Builds are broken https://status.cerebrium.ai/incident/917975 Mon, 08 Jun 2026 13:15:00 -0000 https://status.cerebrium.ai/incident/917975#96832d538ae5c1444e56b3779beb381e68f4509ea8601cfde3f13bc309b84656 Incident US and EU builds are broken Inference Degraded in US East 1 https://status.cerebrium.ai/incident/907571 Thu, 28 May 2026 22:39:00 -0000 https://status.cerebrium.ai/incident/907571#5c020fc321b11bced34ab2fb1eee4e604f6947c610fa0abeafb3a0015dd7f394 Incident all services are restored. we are continuing to monitor Inference Degraded in US East 1 https://status.cerebrium.ai/incident/907571 Thu, 28 May 2026 22:35:00 -0000 https://status.cerebrium.ai/incident/907571#12d08d3c448d8bc05e820bf93644358e7c7cf6fa815dc5bbd3937a1f23f038a0 Incident The issue has been mitigated, and inference should be working as normal. Inference Degraded in US East 1 https://status.cerebrium.ai/incident/907571 Thu, 28 May 2026 22:10:00 -0000 https://status.cerebrium.ai/incident/907571#67b2fe6066908bfe0e3c52eecfff95e1130436426664798a62dbbfc3b395b82f Incident Inference Degraded in US East 1 Inference Degraded in US East 1 https://status.cerebrium.ai/incident/907571 Thu, 28 May 2026 21:35:00 -0000 https://status.cerebrium.ai/incident/907571#91211c5dee1b788bbd39b5b80473f21ce422ce3c6daa4ef990c340dbe4468ea1 Incident Requests are slow and some are failing Unable to schedule workloads in crusoe us-east-1a https://status.cerebrium.ai/incident/900559 Wed, 20 May 2026 19:02:00 -0000 https://status.cerebrium.ai/incident/900559#3dff7911d99d64b4b78b2f5dc0f20774ad55950861341926425d7fe713843936 Incident Upstream provider is back online and all services have been restored. Unable to schedule workloads in crusoe us-east-1a https://status.cerebrium.ai/incident/900559 Wed, 20 May 2026 09:42:00 -0000 https://status.cerebrium.ai/incident/900559#4c1190bc6c91e966bc6ec2789e80910d81dbb0ca696e42827cec5b953ab69ad9 Incident Crusoe is making progress on mitigation in us-east-1. A number of VMs have returned to service, though a subset of compute hosts is still impacted and the incident isn't fully resolved yet. Engineering on the upstream side remains actively engaged. We'll continue posting updates as recovery progresses. Unable to schedule workloads in crusoe us-east-1a https://status.cerebrium.ai/incident/900559 Wed, 20 May 2026 08:27:00 -0000 https://status.cerebrium.ai/incident/900559#ac812de5c081c2a0e0129132ce60f3a912f45791096cf9228c9c7241ae1d7293 Incident Crusoe has a mitigation plan in place and is currently testing it. Early results are positive, with a few more tests to run before applying it across all impacted hosts. We'll post another update once the mitigation is rolled out or if anything changes. Workloads in Crusoe us-east-1 remain affected in the meantime. Unable to schedule workloads in crusoe us-east-1a https://status.cerebrium.ai/incident/900559 Wed, 20 May 2026 03:30:00 -0000 https://status.cerebrium.ai/incident/900559#9bfccd6a7f139d01a8ce3ca57b590ccf039c7f3dc561b203e883562fd4c5ff1a Incident We're seeing a full outage of our Crusoe us-east-1 region affecting all workloads deployed there. The upstream provider is experiencing an infrastructure failure and we're working with them on resolution. Impact: All apps deployed to Crusoe us-east-1 are unavailable. Requests to affected deployments will fail. Workaround: If your app is configured for multi-region or has a fallback region, traffic should route automatically. Customers running solely in Crusoe us-east-1 can redeploy to another region (e.g. AWS us-east-1, AWS us-east-2) in the meantime. We'll post updates here as we have them. Apologies for the disruption. Unable to schedule workloads on Crusoe (both regions) https://status.cerebrium.ai/incident/895827 Thu, 14 May 2026 13:01:00 -0000 https://status.cerebrium.ai/incident/895827#8c0b7508f727bc7304a218e924de8becfee7c1a3e0c765db9526de7c917fa699 Incident Workloads in both Crusoe regions are running normally again. After rolling back the change, we manually cleared the broken state in our infrastructure and brought capacity back online. A post-mortem will follow. Thanks for your patience. Unable to schedule workloads on Crusoe (both regions) https://status.cerebrium.ai/incident/895827 Thu, 14 May 2026 11:49:00 -0000 https://status.cerebrium.ai/incident/895827#171167a905febbe4b8b353b846a7c72b605c33bf81a8dc5e8584576e02c8d3f9 Incident We've identified and mitigated the underlying issue and are now working to bring services back online in both Crusoe regions. Workloads are not yet schedulable while recovery is underway. Next update in 30 minutes or when resolved. Unable to schedule workloads on Crusoe (both regions) https://status.cerebrium.ai/incident/895827 Thu, 14 May 2026 11:07:00 -0000 https://status.cerebrium.ai/incident/895827#6138c3f1bedb7fbbaf6f1409304a8682c9fa60ed882333aa5a727be7ff3c4ef9 Incident We've identified an issue affecting both new deployments and running workloads in our Crusoe regions. A recent internal release is preventing workloads from running correctly. We're rolling back the change now and expect things to recover shortly. Next update in 30 minutes or when resolved. New builds broken in US East https://status.cerebrium.ai/incident/895373 Wed, 13 May 2026 20:45:00 -0000 https://status.cerebrium.ai/incident/895373#9fef88b618a0d2045940e8bde6729c140b953d0e5cbd5cbe65ecd4df0f9afde6 Incident Builds are currently not working in US East Issue routing inference calls to newly deployed applications https://status.cerebrium.ai/incident/884315 Thu, 30 Apr 2026 01:30:00 -0000 https://status.cerebrium.ai/incident/884315#a2ed6f442611bea58d4a5819e72b08f82553fd568b12ac9cf1aee0f6fefad645 Incident Newly deployed apps are responding 404 despite being deployed. Maintenance: Filesystem & Infrastructure Scaling Improvements https://status.cerebrium.ai/maintenance/851226 Tue, 17 Mar 2026 11:34:16 -0000 https://status.cerebrium.ai/incident/851226#6f28296c463fa72e42d355ff8c2e520f3c380aa4c8e13684c2766a1954a50a2d Maintenance We are carrying out planned infrastructure work on 22 March (09:00 EST, for 1 hour) to improve workload scaling and filesystem performance across our clusters. During this window you should expect intermittent downtime that may affect active runs, along with increased latency on deployments and inference requests. Not all regions will be affected simultaneously. US-east-1 is down https://status.cerebrium.ai/incident/843719 Sat, 07 Mar 2026 19:20:00 -0000 https://status.cerebrium.ai/incident/843719#2e1739e0fd9e25dc8dba63407c6472b0c8ffe9acabd0df76baf3ff04b3fa9a1d Incident Our AWS us-east-1 region is down - we have identified the issue and the team is working on resolving it. It has been resolved Increase in 502 errors https://status.cerebrium.ai/incident/819123 Thu, 05 Feb 2026 08:48:00 -0000 https://status.cerebrium.ai/incident/819123#1a6978205cac36d53e7bc13591a56c5a4dc24233b09559ec76f84b222312d822 Incident Some customers are experiencing an increase in 502 errors in US-EAST-1 due to a contention issue on the platform. The team is currently investigating and will revert back as soon as there is more information. We sincerely apologise for this issue and are working to get it resolved as quickly as possible CLI authentication failing https://status.cerebrium.ai/incident/818383 Wed, 04 Feb 2026 10:38:00 -0000 https://status.cerebrium.ai/incident/818383#e00a3f355dc663548b7d3e02bf6b1f26062fb86953948827592deadcb0b42a3f Incident The issue has been resolved. An incorrectly configured DNS record caused users to be unable to sign in using the CLI Increase in request queuing on AWS workloads https://status.cerebrium.ai/incident/802912 Mon, 12 Jan 2026 09:00:00 -0000 https://status.cerebrium.ai/incident/802912#7ab384d4aba5d5dcf8222a210d787a36433e8c2ff3171bf5f85060e21b8cd863 Incident We're currently experiencing degraded performance on workloads being scheduled to the AWS provider. This issue currently only affects GPU-based workloads. This issue is intermittent and may not be affecting all apps. The team is currently investigating the issue and we will provide an update as we uncover any new information. Problem starting new workloads. Existing apps are unaffected. https://status.cerebrium.ai/incident/783164 Tue, 09 Dec 2025 19:08:00 -0000 https://status.cerebrium.ai/incident/783164#3a148ec3d4aca0a7568662b6de72a63d3307b7004630d8db4f39a2d78be6ec4c Incident The issue has been resolved Problem starting new workloads. Existing apps are unaffected. https://status.cerebrium.ai/incident/783164 Tue, 09 Dec 2025 18:44:00 -0000 https://status.cerebrium.ai/incident/783164#830358f004199aa5af28e313f89f76798f7c9008f45ffd0d748217510683a6ce Incident New apps are unable to start at present. Elevated Errors in US-East-1 https://status.cerebrium.ai/incident/778505 Tue, 02 Dec 2025 23:54:00 -0000 https://status.cerebrium.ai/incident/778505#9a6a4a594b4a98c27a6518f481d7a24a1c5d001b1b7369a32cd3ff823a3829aa Incident Our platform is current struggling to schedule new containers on incoming requests. Our team is working on identifying the error and resolving ASAP Resolved: The issue was caused by a failure in a managed component from one of our infrastructure providers, which temporarily prevented us from scheduling new capacity. We’ve worked with the provider to restore functionality and are now implementing additional safeguards to ensure this does not recur. Maintenance: Updating various cluster components https://status.cerebrium.ai/maintenance/765784 Sat, 15 Nov 2025 15:40:01 -0000 https://status.cerebrium.ai/incident/765784#44c162f4b64153670bac6f17c25bfa4e676dc9f436b6a01c2f3a84cc52e0defd Maintenance We are performing a series of infrastructure optimizations to improve performance and reliability. While we don’t expect customer traffic to be impacted, there may be brief periods of elevated latency or volatility during the upgrade window. Our team is closely monitoring the rollout and will update this page with any relevant changes. Maintenance: Emergency node maintenance in US-East-1 https://status.cerebrium.ai/maintenance/757186 Mon, 03 Nov 2025 21:45:27 -0000 https://status.cerebrium.ai/incident/757186#d47ae91f32582e55a5a2dcc9e6bc40e24a2191052cb85532b3e4de37ecdcefe7 Maintenance A critical error in the mechanism GPU devices use to attach to containers is affecting several workloads on the platform, causing NVML to show "Device not found" when calling nvidia-smi or attempting to use the GPU (Mentioned in https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/troubleshooting.html#containers-losing-access-to-gpus-with-error-failed-to-initialize-nvml-unknown-error). This maintenance will update all GPU nodes to use the CDI, as well as a few container runtime upgrades. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 23:36:00 -0000 https://status.cerebrium.ai/incident/746816#def6d05d3ec66619875cc72f480e59b5e4fc16b651f34fefe86783f281a574ad Incident Resolved Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 19:17:00 -0000 https://status.cerebrium.ai/incident/746816#c2796bb6bcf9aeb2ead994d0816196a44f82c71df6e0c874a4db8826faec0b59 Incident We continue to observe recovery across all AWS services, and instance launches are succeeding across multiple Availability Zones in the US-EAST-1 Regions Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 18:24:00 -0000 https://status.cerebrium.ai/incident/746816#08caa7295d5150a3985a744fff62124415b3fe892f24c413ded26bc73486cbcb Incident AWS's mitigations to resolve launch failures for new EC2 instances continue to progress and we are seeing increased launches of new EC2 instances. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 17:48:00 -0000 https://status.cerebrium.ai/incident/746816#63d5cc546455aa540f4c11553e1ee571569501a82cc40bf8190db9cf776ad430 Incident AWS have resolved launch failures and are rolling out the changes to all AZ's at which point we expect launch errors and network connectivity issues to subside. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 17:04:00 -0000 https://status.cerebrium.ai/incident/746816#0344617f774b7af62a0b35ad079fe58cd65549059c279b494e519764a530a924 Incident AWS is in the process of validating a fix for EC2 launches and will deploy to the first AZ as soon as they have confidence we can do so safely. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 15:47:00 -0000 https://status.cerebrium.ai/incident/746816#869f54ebbeff72d23a7f83e3ee9b40b543149323c691df5acef5f39efa5e3be7 Incident AWS have narrowed down the source of the network connectivity issues that have impacted their services. They are throttling requests for new EC2 instance launches to aid recovery and actively working on mitigations. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 14:01:00 -0000 https://status.cerebrium.ai/incident/746816#55398d34ca052b66a663b8a6fafb6229c9d65baf791efe2ae6318d5cc992ecff Incident AWS has applied fixes but is still experiencing problems launching instances in us-east-1. Builds and endpoint calls remain broken. We'll keep you posted. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 13:28:00 -0000 https://status.cerebrium.ai/incident/746816#67599f05df32c991402005470e8eaf57294cf54e0e8c0e1a09a50c5bef88da37 Incident The AWS outage is ongoing. Builds are currently broken due to an outage with EC2. We're waiting on AWS to resolve the issue and will keep you updated. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 11:10:00 -0000 https://status.cerebrium.ai/incident/746816#a49214b9a48be3601aad264a1fdf6dc91ff8867170cd7b4c97618fc61a65bc16 Incident All services have now been restored fully. We will continue to monitor for any anomalies. Thank you for your patience and we apologise for the inconvenience. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 09:43:00 -0000 https://status.cerebrium.ai/incident/746816#15bc346de987b0c270ff70ae21f1a5339045ee3e609949ccc47943dbc02a18d0 Incident Most services have now recovered. You may still experience issues building apps on Cerebrium while AWS continues to resolve the remaining problems. We'll update you once everything is back to normal. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 09:31:00 -0000 https://status.cerebrium.ai/incident/746816#f2755af8f9d9beb9133349e36c2cb6dd9b14b1d56cc67d2fc1b92ca5cee1077f Incident AWS has applied a fix and some services are starting to recover. You may still see some errors or slower response times as things fully stabilize. If something fails, please try again. We'll keep you posted as more services are restored. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 09:01:00 -0000 https://status.cerebrium.ai/incident/746816#7f5682cdb78d1f389b7f350a2e1e75fd236a69267522cbc2fbf643b51989e0ad Incident AWS has identified the root cause as a DNS resolution issue affecting DynamoDB and other services in US-EAST-1. They're working on multiple recovery paths to accelerate the fix. Cerebrium services remain impacted during this time. If you encounter errors, please continue to retry your requests. AWS will provide their next update by 2:45 AM. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 08:29:00 -0000 https://status.cerebrium.ai/incident/746816#871cd301c7bebebef8f179e43876babbef14a6d7fbd37f53a467067d6240c74e Incident The AWS team have narrowed critically affected services down, however, these services are core to the Cerebrium platform and your dashboards, builds, and endpoint calls are still affected. We are continuing to investigate and will provide more updates within the next 45 minutes. Elevated upstream errors (us-east-1) https://status.cerebrium.ai/incident/746816 Mon, 20 Oct 2025 07:38:00 -0000 https://status.cerebrium.ai/incident/746816#5a509bc68dcfde22169faca0750514fa7e5c34b578ec1b50a44df545757ed329 Incident We are seeing elevated error rates from upstream AWS errors across the majority of our services in the us-east-1 region. We will share an update as soon as possible. Degraded Inference API in US-EAST-1 https://status.cerebrium.ai/incident/740083 Wed, 08 Oct 2025 18:28:00 -0000 https://status.cerebrium.ai/incident/740083#89d8d4e8dd689746d3c782842aa817ffbafe52467e5adfe5607a9365aceac920 Incident The Inference API is currently experiencing degraded performance in US-EAST-1. Our team is working on a fix ASAP Inference API https://status.cerebrium.ai/incident/737024 Fri, 03 Oct 2025 13:13:00 -0000 https://status.cerebrium.ai/incident/737024#e1c9c6e4c4fdf5e170832cdabbc8311af2f5a5ebda3688b6218f48be2e12c17e Incident Inference API is currently experiencing a High 502 failure rate. Roughly 45% of all requests are affected. Our team is currently investigating the cause of the issue as a matter of high urgency. Container Count is down https://status.cerebrium.ai/incident/726877 Thu, 18 Sep 2025 23:05:00 -0000 https://status.cerebrium.ai/incident/726877#02b232bc3e6fa758f3a2ce6d5b1043c6c6f3e30c2573ee0fa7cd895a0e38bb5f Incident A 3rd party provider is down affecting the container count on the dashboard.