Intelligence
AI Evaluation Engineer
Decide what done means. If we cannot score it, we have not specified it, and we certainly cannot sell it.
- Location
- Lahore or remote
- Contract
- Permanent, full time
- Level
- Mid to senior
- Compensation
- Competitive, reviewed against market twice a year
About the role
We sell fixed-scope work, which means we need a defensible answer to whether a thing is finished. For deterministic software that is a test suite. For a system that is probabilistic on purpose it is an evaluation harness, and building good ones is a genuine speciality. This role owns that speciality across every engagement and both products.
What you will do
- Turn client requirements into measurable evaluation sets, including the adversarial cases nobody asked for.
- Build the harnesses and the tooling around them so any engineer can run and read them.
- Track drift in production and raise it before a client does.
- Work on BotUp, where the evaluation harness is not internal tooling but the product itself.
- Push back on scope that cannot be measured, early, while it is still cheap to change.
What we need from you
- You have built evaluation for a machine learning or LLM system that other people then relied on.
- Statistical literacy: you know what a small sample can and cannot tell you.
- Strong Python and a taste for tooling that other engineers actually adopt.
- Scepticism. This role only works if you are willing to report a result nobody wants.
Useful but not required
- Testing or QA engineering background before moving into ML.
- Experience with human-in-the-loop annotation at any scale.
- You have written about evaluation somewhere public.
Apply
Apply for AI Evaluation Engineer
This goes straight to the engineering team. We reply either way inside two weeks, and if it is a no you get told why.
Other openings








