Senior Infrastructure & Operations Engineer – Kubernetes / Platform Reliability
London (hybrid, 2-3 days/week in Paddington)
We’re working with a well-funded, fast-scaling treasury firm managing over $30bn in assets, hiring a senior infrastructure engineer to own platform reliability.
The core of the role is deep, hands‑on Kubernetes – cluster internals, networking, storage and security – across a bare‑metal and cloud estate.
You’ll design, operate and harden production clusters, build the monitoring and alerting the platform runs on, and keep services stable under real load and failure.
What you’ll own:
- Production Kubernetes end to end: cluster builds, upgrades, hardening, performance – on bare-metal and AWS.
- The observability stack (metrics, logging, tracing, alerting) and the SLOs/SLIs behind it.
- The edge: Cloudflare caching, routing and security rules.
- CI/CD that’s fast, reproducible and safe – rollouts, canaries, rollbacks – with everything as IaC.
- Incidents and postmortems, and designing out single points of failure.
What you’ll need:
- Senior-level experience running production infrastructure, with strong Linux fundamentals.
- Strong networking foundations – TCP/IP, routing, DNS, TLS, load balancing – and a record of debugging distributed‑system faults.
- Experience building monitoring and observability systems with reliable, low‑noise signals.
Bonus: multi‑cluster/multi‑region setups, service meshes, Elasticsearch or other large‑scale data infrastructure.
Comp: Up to £125k base plus meaningful equity.
If this sounds like you, please apply and we’ll reach out with more information.
#J-18808-Ljbffr…
