Salary: £70,000 – 105,000 per year
Requirements
- Strong backend/systems engineering background with experience building and operating production services with reliability and observability requirements.
- Experience designing and delivering shared platform or infrastructure components used by multiple teams.
- Strong production ownership, including monitoring, alerting, incident response, debugging, and post-incident learning.
- Distributed systems fundamentals, including async workflows, idempotency, consistency trade-offs, and designing for failure.
- Hands‑on experience with LLM APIs or a strong interest in learning their production failure modes, such as rate limits, context windows, multi‑vendor routing, latency variance, and cost control.
- Security mindset for AI systems, including prompt injection risks, PII in logs, data leakage, and safe credential handling.
- Strong programming experience in either a JVM‑based language or Python; we operate a polyglot platform with components written in Kotlin and Python, and youll be expected to contribute to both.
- Clear communication and collaboration skills, especially when working with product teams and other engineers to turn ambiguous platform needs into practical solutions.
- Passion for developer experience and a desire to be a force multiplier for engineering teams.
- Deep understanding of the realities of LLMOps, data retrieval, prompt and context engineering, and model evaluation in production.
- Comfort switching languages or technologies to achieve goals.
- Proven experience building platforms or tooling for agentic AI.
- Ability to work fully autonomously on an entire product feature from design to implementation.
Responsibilities
- Design, build, and operate core GenAI platform components used by product teams at Pleo, including the LLM routing gateway, vector search and RAG infrastructure, tool registry and MCP gateway, AI observability and evaluation tooling, and infrastructure for multi‑step, long‑running agentic workflows.
- Own production‑quality delivery of platform features from design through rollout, monitoring, and follow‑up.
- Contribute to resilient system design, including sensible APIs, failure handling, rate limiting, retries, idempotency, and safe change management.
- Improve reliability and observability through metrics, dashboards, alerting, incident follow‑ups, and operational improvements.
- Partner with Applied AI Engineers and product teams to understand platform needs and help them build AI‑powered features safely.
- Build internal SDKs, templates, and guardrails that help product engineers build AI features without deep infrastructure expertise.
- Support other engineers through pairing, code reviews, technical feedback, and clear documentation.
- Help evaluate build‑vs‑buy decisions in the rapidly evolving LLMOps tooling landscape.
- Develop a clear picture of how AI features are currently being built at Pleo and where the biggest infrastructure bottlenecks are.
- Take ownership of a core platform component and improve its reliability, observability, or developer experience.
- Deliver production‑ready improvements with clear rollout plans, monitoring, and operational documentation.
- Partner with Applied AI Engineers and product teams to identify platform investments that would accelerate their work.
- Contribute to our internal standards for AI feature development, including how we evaluate quality, manage prompts, and monitor production AI systems.
Technologies
- AI
- Backend
- Support
- JVM
- Kotlin
- LLM
- Python
- Security
- Windows
About the role
We are Pleo, a spend management company on a mission to make managing money seamless, empowering, and effective for finance teams and employees. We build AI‑powered and other spend solutions with a vision to help businesses go beyond, and we are a driven, progressive, kind team of 850+ people from over 100 nationalities serving 40,000+ customers. This role sits in our GenAI Core team, where we build the horizontal platform infrastructure behind our AI features, with close collaboration across Engineering, Applied AI, Data Science, and product teams. We offer remote, hybrid, or in‑person working in eligible locations, subject to valid right to work in the country of choice and no visa sponsorship. Our benefits include a Pleo card, lunch support, private healthcare, 25 days of holiday plus public holidays, hybrid and remote options, MyndUp mental health support, and paid parental leave.
#J-18808-Ljbffr…
