About this role
About the Role Rootstrap is seeking a Sr. QA/Eval Engineer to test AI-powered applications, with a focus on evaluating non-deterministic AI systems (LLMs, AI agents, and RAG applications) and building or maintaining evaluation frameworks to measure quality, detect regressions, and improve AI performance. What You'll Do
- Create detailed, comprehensive, and well-structured test plans and test cases.
- Review client needs and technical specifications with the development team to provide timely and meaningful feedback.
- Estimate, prioritize, plan, and coordinate testing activities.
- Identify, record, document thoroughly, and track bugs.
- Design, develop, and execute automation scripts (including Playwright-based automation).
- Perform thorough regression testing when bugs are resolved.
- Track quality assurance metrics and reports.
- Design and maintain evaluation suites (evals) to validate AI-powered features and detect regressions in model behavior.
- Test non-deterministic AI systems, including LLM-based features, AI agents, and RAG applications, using qualitative and quantitative evaluation methods.
- Analyze AI outputs, identify failure patterns (hallucinations, instruction-following issues, grounding, etc.), and collaborate with engineers to improve system quality.
- Stay up to date with new testing tools, AI evaluation techniques, and quality assurance strategies while driving continuous improvement.
- Investigate the causes of software defects and collaborate with the development team to implement preventive solutions.
- Collaborate on internal initiatives. What We're Looking For
- Background in Software Engineering, Computer Science, or a related field.
- 3+ years of experience as a QA Engineer.
- Advanced English skills.
- Hands-on experience with mobile testing (required).
- Experience with test automation, preferably using Playwright.
- Experience using AI tools (e.g., Claude or similar) to optimize testing processes, including test case creation, bug reporting, edge case identification, and automation support.
- Experience testing AI-powered applications or LLM-based systems.
- Understanding of the challenges of testing non-deterministic systems and validating AI-generated outputs.
- Experience with evaluation engineering, including building or maintaining evals, benchmark datasets, or automated AI quality evaluations. Nice to Have
- Knowledge of CI/CD pipelines and automated test integration.
- Experience testing AI agents, RAG applications, or conversational AI systems.
- Familiarity with AI evaluation frameworks or LLM observability tools such as LangSmith, Braintrust, Arize Phoenix, Promptfoo, OpenEvals, or similar. Compensation & Benefits
- Access to Rootstrap University, Conferences/Certifications, and a Mentorship Program for your professional growth.
- Learning Bonus.
- Opportunities to organize cross-functional initiatives and receive 360 Feedback to improve your skills continuously.
- The flexibility to work remotely or from our offices in Montevideo, Buenos Aires, and Medellín, with a flexible time schedule and Workation program.
- Gym benefits, psychological counseling, and weekly lunch reimbursements with special foods in the offices.
- An Onboarding kit and access to cutting-edge technologies and tools to make your work easier.
- After offices, Prizes, and special occasions gifts to celebrate your achievements.
- People Care referent to support your well-being and personal development.
- We value your well-being and happiness. We grow together!