Talent Apply
Log in
All jobs
R

Sr. QA/Eval Engineer

rootstrap
Remote (Argentina)
Remote

About this role

About the Role Rootstrap is seeking a Sr. QA/Eval Engineer to test AI-powered applications, with a focus on evaluating non-deterministic AI systems (LLMs, AI agents, and RAG applications) and building or maintaining evaluation frameworks to measure quality, detect regressions, and improve AI performance. What You'll Do

  • Create detailed, comprehensive, and well-structured test plans and test cases.
  • Review client needs and technical specifications with the development team to provide timely and meaningful feedback.
  • Estimate, prioritize, plan, and coordinate testing activities.
  • Identify, record, document thoroughly, and track bugs.
  • Design, develop, and execute automation scripts (including Playwright-based automation).
  • Perform thorough regression testing when bugs are resolved.
  • Track quality assurance metrics and reports.
  • Design and maintain evaluation suites (evals) to validate AI-powered features and detect regressions in model behavior.
  • Test non-deterministic AI systems, including LLM-based features, AI agents, and RAG applications, using qualitative and quantitative evaluation methods.
  • Analyze AI outputs, identify failure patterns (hallucinations, instruction-following issues, grounding, etc.), and collaborate with engineers to improve system quality.
  • Stay up to date with new testing tools, AI evaluation techniques, and quality assurance strategies while driving continuous improvement.
  • Investigate the causes of software defects and collaborate with the development team to implement preventive solutions.
  • Collaborate on internal initiatives. What We're Looking For
  • Background in Software Engineering, Computer Science, or a related field.
  • 3+ years of experience as a QA Engineer.
  • Advanced English skills.
  • Hands-on experience with mobile testing (required).
  • Experience with test automation, preferably using Playwright.
  • Experience using AI tools (e.g., Claude or similar) to optimize testing processes, including test case creation, bug reporting, edge case identification, and automation support.
  • Experience testing AI-powered applications or LLM-based systems.
  • Understanding of the challenges of testing non-deterministic systems and validating AI-generated outputs.
  • Experience with evaluation engineering, including building or maintaining evals, benchmark datasets, or automated AI quality evaluations. Nice to Have
  • Knowledge of CI/CD pipelines and automated test integration.
  • Experience testing AI agents, RAG applications, or conversational AI systems.
  • Familiarity with AI evaluation frameworks or LLM observability tools such as LangSmith, Braintrust, Arize Phoenix, Promptfoo, OpenEvals, or similar. Compensation & Benefits
  • Access to Rootstrap University, Conferences/Certifications, and a Mentorship Program for your professional growth.
  • Learning Bonus.
  • Opportunities to organize cross-functional initiatives and receive 360 Feedback to improve your skills continuously.
  • The flexibility to work remotely or from our offices in Montevideo, Buenos Aires, and Medellín, with a flexible time schedule and Workation program.
  • Gym benefits, psychological counseling, and weekly lunch reimbursements with special foods in the offices.
  • An Onboarding kit and access to cutting-edge technologies and tools to make your work easier.
  • After offices, Prizes, and special occasions gifts to celebrate your achievements.
  • People Care referent to support your well-being and personal development.
  • We value your well-being and happiness. We grow together!

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →