We build AI-powered solutions including conversational AI, RAG systems, AI agents, and enterprise automation platforms.RoleJoversational AI, and RAG applications. Focus on AI quality, benchmarking, safety, observability, and automation to ensure production-ready systems.ResponsibilitiesTest Generative AI and conversational systemsPerform LLM evaluation and benchmark testingValidate RAG (retrieval quality, relevance, grounding)Build automation frameworks (Playwright/Selenium)Conduct API, regression, and E2E testingPerform AI safety, red teaming, and bias validationSupport observability, monitoring, and HITL workflowsCollaborate with engineering, QA, and DevOps teamsSupport CI/CD quality pipelinesRequired SkillsQA & automation testingPlaywright/Selenium, API testing (Postman, REST)GenAI & RAG validation, LLM evaluationPrompt engineering, conversational AI testingHITL testing, AI safety & observabilitySQL, Git, Python, CI/CD conceptsPreferredOpenAI, Azure OpenAI, ChatGPT, ClaudeLangChain, CrewAI, LangGraph, MCPVector DBs (Pinecone, ChromaDB, Weaviate)RAGAS, DeepEvalPython/JS, Docker, Kubernetes, GitHub Actions ( US Remote 5 position )