AI Evaluation Engineer

Company: Gazelle Global
Apply for the AI Evaluation Engineer
Location: London
Job Description:

Location: London (Hybrid – 1 day onsite)

Contract: 6–12 Months

We’re looking for an AI Evaluation Engineer to help deliver and optimise cutting‑edge Generative AI solutions within a large enterprise environment.

Key Responsibilities

  • Define and implement evaluation strategies for LLMs, RAG pipelines and AI agents
  • Build automated evaluation pipelines and benchmarking frameworks
  • Establish evaluation metrics and create test datasets
  • Evaluate prompt quality, model performance and response accuracy
  • Validate RAG knowledge grounding and conduct safety, risk and compliance testing
  • Support human‑in‑the‑loop evaluation and continuous AI optimisation

Required Skills

  • Strong experience with LLMs, RAG architectures and prompt engineering
  • Python for data analysis and evaluation pipelines
  • Experience with AI evaluation tools such as DeepEval, PromptTools or similar
  • Knowledge of NLP quality assessment, benchmarking and experimentation frameworks
  • Understanding of Responsible AI principles

#J-18808-Ljbffr…

Posted: July 18th, 2026