Resolved -
A production Kubernetes node in the cloud provider failed and repeatedly dropped offline, which took the ingress gateway and several backend services with it. We added capacity, removed the failing node, and moved all workloads to healthy nodes.
Sep 2, 20:57 EDT
Monitoring -
As of 00:12 PM UTC (8:12PM ET), ViiBE services have restarted and the portals are loading again. A production Kubernetes node in the cloud provider failed and repeatedly dropped offline, which took the ingress gateway and several backend services with it. We added capacity, removed the failing node, and moved all workloads to healthy nodes. We are monitoring closely while services finish recovering.
Sep 2, 20:15 EDT
Update -
We are continuing to investigate this issue.
Sep 2, 19:49 EDT
Investigating -
Starting at approximately 22:46 UTC (6:46 PM ET) on September 2, ViiBE portals are intermittently unreachable. Users may see connection failures or a "no healthy upstream" error. The cause is a production Kubernetes node that is repeatedly going unhealthy, which is disrupting the ingress gateway that serves these sites. We have identified the failing node and are working to move traffic off it. Next update in 30 minutes.
Sep 2, 19:48 EDT