AI Conversational Engineer
Conversational AI fails in ways that are hard to see until a user hits them: intent mismatches, context collapse mid-dialogue, hallucinated responses that sound plausible. This role exists to catch those failures before they reach production — and to build the tooling and test strategy that keeps catching them as models evolve.
You will work across NLP pipelines, LLM integrations and full conversational workflows, validating not just whether a response comes back but whether it is accurate, contextually appropriate and usable. You will define the KPIs — intent accuracy, response latency, conversation success rate — that give the team an honest signal on quality.
Who thrives here
You have spent at least three years in QA, ideally on AI, ML or NLP products, and you understand what makes conversational systems genuinely difficult to test: state, context, non-determinism and the long tail of unexpected inputs. You are comfortable in code — Python, JavaScript or similar — and you know your way around API testing and CI/CD pipelines. You are also expected to be AI-native: fluent in AI tooling, and curious about the applications reshaping your domain.
What you'll own
- Develop and execute test plans covering NLP pipelines, dialogue flows, integrations and APIs
- Design functional, regression, performance and exploratory tests for conversational interfaces
- Validate intent recognition, entity extraction, context handling and multi-turn dialogue behaviour
- Test conversational flows for accuracy, completeness and user experience, including edge-case and unexpected-input scenarios
- Build automated tests for API endpoints, dialogue flows and ML outputs using Postman, Cypress, PyTest or custom NLP tooling
- Contribute automated regression suites to CI/CD pipelines
- Partner with Data Scientists, ML Engineers and PMs on acceptance criteria and conversational design feedback
- Define and track quality KPIs — intent accuracy, response latency, conversation success rate
- Uphold accuracy, robustness, fairness and usability standards across all conversational systems
What it takes
- 3+ years of QA/testing experience, preferably with AI, ML or NLP products
- Strong API testing skills, automated framework experience, and scripting in Python, JavaScript or similar
- Familiarity with NLP concepts — intents, entities, context, LLMs and dialogue frameworks such as Rasa, Dialogflow or LLM APIs
- Experience with CI/CD pipelines and version control (Git)
- AI-native fluency: active use of AI tools and working knowledge of AI applications in the domain
- Direct experience testing LLM-based chat or voice agents, including prompt-response validation and hallucination detection
- Knowledge of model-evaluation techniques — A/B testing, offline eval sets and human-in-the-loop scoring
- Understanding of AI ethics, fairness and bias testing for conversational systems
- Experience with accessibility and localisation testing for multilingual conversational agents
Nice to have
- Hands-on A/B testing and model evaluation experience
- Cloud platform experience — AWS, GCP or Azure
- Accessibility and localisation work for multilingual agents
- Active interest in AI ethics, fairness and bias testing
We respond to every application within 2 business days.