Nexaminds Logo

Nexaminds

Senior Quality Engineer (AI)

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in Canada
Senior level
Remote
Hiring Remotely in Canada
Senior level
Lead AI validation strategy for LLMs, RAG, and agent systems. Build automated CI/CD validation pipelines, design semantic and prompt regression tests, implement guardrails and confidence scoring, monitor production AI quality, and partner with engineering and data governance to ensure safe, accurate, production-ready AI.
The summary above was generated by AI

Unlock Your Future with Nexaminds!

At Nexaminds, we're on a mission to redefine industries with AI. We're passionate about the limitless potential of artificial intelligence to transform businesses, streamline processes, and drive growth.

Join us on our visionary journey. We're leading the way in AI solutions, and we're committed to innovation, collaboration, and ethical practices. Become a part of our team and shape the future powered by intelligent machines. If you're driven by ambition, success, fun, and learning, Nexaminds is where you belong.

Nexaminds is looking for a Senior Quality Engineer (AI) to lead the validation and quality strategy for AI-powered systems, including Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) applications, and agent-based workflows. The ideal candidate has deep experience in Quality Engineering, AI evaluation, and automated validation frameworks, with a strong understanding of how to measure, monitor, and improve the reliability of non-deterministic AI systems. This role focuses on building enterprise-grade AI validation capabilities, designing automated evaluation pipelines, implementing AI guardrails, and partnering closely with engineering, AI/ML, and data governance teams to ensure AI-powered solutions are accurate, safe, and production-ready.

Location: Canada (Remote)

Qualifications we are looking for:

  • 7+ years of experience in Quality Engineering, Software Development, or a related technical discipline, including recent hands-on experience with AI/ML or LLM-based systems.
  • Experience designing validation or evaluation frameworks for non-deterministic or probabilistic systems such as Large Language Models (LLMs) or Machine Learning applications.
  • Strong understanding of Retrieval-Augmented Generation (RAG), agentic workflows, and LLM evaluation concepts, including hallucination detection, groundedness, relevance, and consistency.
  • Hands-on experience with Large Language Models (LLMs) and prompt engineering, including prompt design, optimization, and regression validation.
  • Practical experience using Claude Code, Claude SDK, or similar AI-assisted development platforms and agentic coding tools.
  • Proficiency with Node.js for building automation tools, validation pipelines, and engineering utilities.
  • Experience working with MongoDB to store, manage, and analyze evaluation results, validation data, or AI output logs.
  • Experience integrating automated validation processes into CI/CD pipelines.
  • Experience designing semantic validation approaches that evaluate meaning, relevance, and contextual accuracy beyond traditional software testing.
  • Strong understanding of AI quality metrics, confidence scoring, output consistency, and production monitoring.
  • Solid knowledge of data validation principles, including schema validation, business rule compliance, and data consistency.
  • Excellent cross-functional collaboration skills, with experience partnering across Quality Engineering, AI/ML Engineering, Software Engineering, and Data Governance teams.

Nice to have:

  • Experience with prompt regression testing tools or frameworks.
  • Familiarity with AI governance, fairness, or bias-detection practices.
  • Experience with tools such as Playwright, RestSharp, or similar UI/API test automation.
  • Prior experience standing up a validation or evaluation function from scratch.
  • Exposure to confidence scoring or guardrail systems for production AI.
  • Experience designing or orchestrating multi-agent systems in production environments

Job duties:

  • Define and lead the AI validation strategy for LLMs, RAG systems, and AI agents.
  • Build automated validation pipelines integrated into CI/CD.
  • Establish evaluation frameworks covering correctness, relevance, groundedness, consistency, and hallucination rate.
  • Design and maintain prompt regression testing to catch silent quality degradation after model or prompt changes.
  • Introduce semantic validation techniques that go beyond traditional QA (evaluating meaning, not just structure).
  • Lead production monitoring of AI output quality and define acceptable output ranges for non-deterministic systems.
  • Partner with engineering teams to validate AI-generated code, test cases, and AI-assisted decision workflows.
  • Introduce AI guardrails and confidence scoring to flag unsafe or low-confidence outputs.
  • Build validation checkpoints across multi-step agent pipelines.
  • Define and track quality metrics/KPIs (defect reduction, output consistency, prompt regression stability, adoption confidence).

What you can expect from us

Here at Nexaminds, we're not your typical workplace. We're all about creating a friendly and trusting environment where you can thrive. Why does this matter? Well, trust and openness lead to better quality, innovation, commitment to getting the job done, efficiency, and cost-effectiveness.

  • Stock options 📈
  • Remote work options 🏠
  • Flexible working hours 🕜
  • Benefits above the law
  • But it's not just about the work; it's about the people too. You'll be collaborating with some seriously awesome IT pros.
  • You'll have access to mentorship and tons of opportunities to learn and level up.

Ready to embark on this journey with us? 🚀🎉 If you're feeling the excitement, go ahead and apply!

Similar Jobs

6 Days Ago
In-Office or Remote
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead architecture and technical strategy for AI-driven product quality systems using LLMs and agents. Build scalable evaluation frameworks, detect regressions, generate insights, and drive cross-functional adoption while mentoring engineers and defining standards for trustworthy AI.
Top Skills: AgentsAi InfrastructureEvaluation SystemsLlmsRetrieval Architectures
14 Days Ago
In-Office or Remote
Senior level
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Design, build, and maintain automated test suites, evaluation scenarios, and CI/CD quality gates for AI-powered applications. Test APIs, validate ML model outputs, create measurable evaluation rubrics, mentor teams, and contribute to shared QA frameworks and codebases.
Top Skills: Ai-Powered Testing ToolsAirflowAWSAzureCadCometGithub ActionsGoogle Cloud PlatformGreat ExpectationsJenkinsMetaflowMlflowPythonPyTorchTensorFlowWeights & Biases
3 Hours Ago
Remote
Expert/Leader
Expert/Leader
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
As a Staff Fullstack Software Engineer, you'll optimize user monetization systems at Dropbox, leading technical strategies for checkout and subscriptions while mentoring engineers.
Top Skills: GoPythonReactTypescript

What you need to know about the Singapore Tech Scene

The digital revolution has driven a constant demand for tech professionals across industries like software development, data analytics and cybersecurity. In Singapore, one of the largest cities in Southeast Asia, the demand for tech talent is so high that the government continues to invest millions into programs designed to develop a talent pipeline directly from universities while also scaling efforts in pre-employment training and mid-career upskilling to expand and elevate its workforce.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account