B
Bruno Arndt
Bruno Arndt
About
Detail
Software Engineer III at Loggi
São Paulo, Brazil
Senior Software Engineer and AI Quality & Evaluation Lead with over a decade of experience building and scaling production systems across AI, SaaS, education, and public-sector platforms. Most recently, I worked with G2i, focusing on AI quality assurance, LLM evaluation, and human-in-the-loop review systems, combining hands-on technical work with leadership responsibilities in distributed teams.
My work at G2i and Outlier centered on ensuring the quality, reliability, and testability of AI systems, particularly for AI-assisted code generation. I reviewed and validated AI training tasks and model outputs for clarity, feasibility, correctness, and edge-case coverage; performed comparative evaluation and response ranking; identified systemic failure patterns; and helped define consistent quality standards across datasets and reviewers. In parallel, I contributed to the design and implementation of automated evaluation pipelines that execute real-world software engineering tasks with an emphasis on reproducibility, determinism, and traceability.
As a Tech Lead and Squad Leader, I mentored and supervised distributed teams of senior contributors, aligning individual output to shared quality criteria through structured coaching, reviewer calibration, clear documentation, and continuous feedback loops. I’m particularly strong at bridging technical depth with quality governance, improving delivery speed without sacrificing rigor.
My technical background includes Python, JavaScript/TypeScript, Node.js, AWS, Cloudflare, and modern web architectures, with prior experience designing backend systems in Django, PHP, PostgreSQL, and GraphQL-based platforms. What drives me is building trustworthy AI systems and scalable evaluation processes, while helping people grow through clear standards, thoughtful feedback, and strong engineering culture.
My work at G2i and Outlier centered on ensuring the quality, reliability, and testability of AI systems, particularly for AI-assisted code generation. I reviewed and validated AI training tasks and model outputs for clarity, feasibility, correctness, and edge-case coverage; performed comparative evaluation and response ranking; identified systemic failure patterns; and helped define consistent quality standards across datasets and reviewers. In parallel, I contributed to the design and implementation of automated evaluation pipelines that execute real-world software engineering tasks with an emphasis on reproducibility, determinism, and traceability.
As a Tech Lead and Squad Leader, I mentored and supervised distributed teams of senior contributors, aligning individual output to shared quality criteria through structured coaching, reviewer calibration, clear documentation, and continuous feedback loops. I’m particularly strong at bridging technical depth with quality governance, improving delivery speed without sacrificing rigor.
My technical background includes Python, JavaScript/TypeScript, Node.js, AWS, Cloudflare, and modern web architectures, with prior experience designing backend systems in Django, PHP, PostgreSQL, and GraphQL-based platforms. What drives me is building trustworthy AI systems and scalable evaluation processes, while helping people grow through clear standards, thoughtful feedback, and strong engineering culture.
Contact Bruno regarding:
work
Full-time jobs