Nebius in London is building a cutting‑edge AI cloud platform. This role owns the reliability, performance and observability of the entire inference stack, with hands‑on work on GPUs, Kubernetes and scalable cloud infrastructure.
You will implement telemetry pipelines (metrics, logs, traces), tune autoscalers with Kubernetes, craft Terraform modules, and harden routing and retries. You’ll drive post‑mortems to prevent recurrence while pushing for cost‑effective, high‑reliability operations that
#J-18808-Ljbffr…
