Position OverviewWe’re looking for a Senior Testing Engineer to lead quality across complex software features, AI/ML systems, and integrated platforms. This role is ideal for an experienced engineer who is knowledgeable across both automation and AI testing, enjoys hands-on work, and is growing into owning testing strategy and architecture.In this role, you will work as part of a cross-functional team, designing automation frameworks, validating ML models and AI workflows, and driving quality throughout the delivery process. You’ll collaborate closely with developers, ML engineers, and product teams to deliver high-quality features and develop the technical instincts that come from shipping things that matter.Why This Role MattersAt Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate.Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production-ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.What You’ll DoCraft & DeliveryDesign and develop scalable automation frameworks for UI, API, and integration testing (e.g., Cypress, Playwright, Selenium, PyTest, TestNG, Cucumber)Integrate automated tests into CI/CD pipelines and drive CI/CD quality gates (e.g., GitHub Actions, GitLab CI, Jenkins)Design AI-specific test strategies and validate ML models across key metrics including accuracy, precision, recall, bias, and driftTest data pipelines, feature engineering processes, and validate LLM responses for hallucination risk and output consistencyPerform test planning, review automation code, and ensure best practices are applied across the test codebaseMonitor model quality in production environments and support release validation and quality assurance processesBring an AI-forward mindset to your daily work, using tools like Claude, Cursor, and other modern AI assistants to ship higher-quality work at paceCollaboration & CommunicationCollaborate closely with developers, ML engineers, DevOps, and product teams across the full SDLCCommunicate quality status, testing findings, and risks clearly to stakeholdersParticipate actively in sprint ceremonies, design reviews, and release planningLeadership & InfluenceLead quality efforts end-to-end with growing ownership of testing strategy and automation architectureContribute to QA standards and best practices, and identify opportunities to improve coverage, reliability, and automationBegin mentoring junior engineers, sharing knowledge and supporting their growthWhat You’ll Bring3–4+ years of experience in QA, data QA, or AI testing, with solid knowledge across both automation and AI testingStrong programming skills in Python and at least one other language (e.g., Java, JavaScript)Hands-on experience designing and maintaining automation frameworks (e.g., Selenium, Cypress, Playwright, PyTest, TestNG, Cucumber)Strong understanding of the ML lifecycle and experience with model evaluation metrics (accuracy, precision, recall, bias, drift)Knowledge of AI testing methodologies including LLM validation, hallucination testing, and data pipeline validationExperience with CI/CD pipelines, API automation, and version control (e.g., GitHub Actions, Postman, Git)Familiarity with MLOps pipelines and model monitoring in production environmentsStrong understanding of Agile methodologies and experience with test management tools (e.g., Jira, TestRail)Demonstrable usage of AI-forward tools such as Claude and CursorExperience with mobile automation, performance testing, LLM testing frameworks, prompt engineering, or cloud ML platforms is a plus (e.g., Appium, JMeter, AWS SageMaker, Azure ML)Helpful Extras and Unique SkillsYou’ll Do Well Here if You AreA doer. You see something broken and fix it. You'd rather move on clarity than wait for certainty.A fast learner who knows you don't know everything. The AI landscape changes weekly. You're senior enough to know better and curious enough to keep learning anyway.Direct in a way that makes the work better. You give honest feedback. You'd rather have the hard conversation than blow smoke.Obsessed with craft. You know genius is in the details. You ship exceptional, not perfect, and you don't put your name on work you wouldn't stand behind.Built for ownership. You honor commitments, admit mistakes fast, and back your teammates when a decision costs something. No handoffs, no finger-pointing.All in. You treat clients' businesses like your own. You take the work seriously without taking yourself seriously.Resourceful when the budget, timeline, or team is tight. Constraints don't slow you down. They sharpen you.Glad to be in the room with people who care as much as you do. Our teams average fifteen-plus years of experience. We hire people who push each other to do better work.