Achuth Reddy Bangaru

Achuth Reddy Bangaru

About

Detail

AI/ML Engineer | MS Computer Science @ UAB | LLM Inference Optimization, RAG, Generative AI, PyTorch | Machine Learning & AI Systems
Birmingham, Alabama, United States

Contact Achuth regarding: 
work
Full-time jobs

Timeline


work
Job
school
Education
folder
Project

Résumé


Jobs verified_user 0% verified
  • University of Alabama at Birmingham
    Machine Learning Research Assistant — NVL Lab
    University of Alabama at Birmingham
    Feb 2026 - Current (7 months)
    Developing machine learning pipelines for multimodal neural–behavioral analysis in the Neural Value Laboratory (NVL Lab), University of Alabama at Birmingham. • Engineered GPU-accelerated multimodal ML pipelines in PyTorch for neural-behavioral data analysis on SLURM-based HPC infrastructure, enabling scalable behavioral state decoding across 60+ neuroscience experiments. • Fused DeepLabCut pose estimation, CEBRA contrastive embeddings, and LSTM sequence modeling to decode temporal behavioral patterns and learn high-dimensional neural representations, achieving ~80% accuracy across behavioral conditions. • Parallelized distributed GPU training and cross-modal synchronization on CUDA-enabled HPC systems, cutting experimental runtime by ~30%
  • Ramp
    Artificial Intelligence Engineer
    Ramp
    May 2025 - Dec 2025 (8 months)
    • Developed an LLM-powered financial assistant using LangChain, Hugging Face Transformers, and LoRA-based fine-tuning on Mistral 7B models to automate expense categorization and finance-related query resolution, reducing task resolution time by 45% across internal workflows. • Shipped a RAG pipeline combining Pinecone hybrid retrieval and cross-encoder reranking over internal financial knowledge sources, improving response relevance by 30% across finance support operations. • Streamlined approval routing, anomaly detection, and transaction classification workflows on GCP using Vertex AI and distributed real-time inference pipelines, reducing manual review effort by 40% while improving classification precision across finance operations. • T
  • Razorpay
    Machine Learning Engineer
    Razorpay
    Jan 2023 - Jul 2024 (1 year 7 months)
    • Trained fraud detection and transaction classification models using XGBoost and Scikit-learn on 200K+ payment transactions, optimized via RandomizedSearchCV and Stratified K-Fold cross-validation, achieving 92% detection accuracy while substantially reducing false-positive transaction blocks. • Launched a low-latency payment risk scoring system with FastAPI and real-time ML inference pipelines, conducting A/B testing against legacy rule-based systems and improving payment legitimacy decision speed by 35%. • Improved payment failure prediction and anomaly detection models from transaction logs, improving payment success rates by 20% while reducing fraud leakage by 18%. • Wired up scalable feature engineering and ML data pipelines using AWS
  • Razorpay
    ML Trainee
    Razorpay
    Jun 2022 - Dec 2022 (7 months)
    • Worked on fraud pattern analysis and transaction data preprocessing pipelines to support model training workflows across Razorpay's payment infrastructure. • Assisted in building and validating early-stage ML models for payment anomaly detection, performing feature engineering and data cleaning on large-scale transaction datasets using Python and Scikit-learn.
Education verified_user 0% verified
  • University of Alabama at Birmingham
    University of Alabama at Birmingham
    University of Alabama at Birmingham
    Jan 2024 - Jul 2026 (2 years 7 months)
    Relevant Coursework: Artifical Intelligence, Machine Learning, Deep Learning, Natural Language Processing, Computer Vision,Foundation of Data Science, Advanced Algorithms and Applications, Cloud Security
  • Karunya Institute of Technology and Sciences
    Karunya Institute of Technology and Sciences
    Karunya Institute of Technology and Sciences
Projects (professional or personal) verified_user 0% verified
  • S
    Systems-Level Optimization of LLM Inference on a Single Consumer GPU
    Dec 2025 - Mar 2026 (4 months)
    Engineered production-grade LLM inference server on a single 8GB NVIDIA RTX GPU — achieving 3× throughput, 74% latency reduction, and 69% cost savings through async batching and KV cache optimization. • Engineered async dynamic batching server (20 ms coalescing window) boosting throughput 3× (0.57 → 1.73 req/s) and reducing p50 latency by 74% (8.3s → 2.2s). • Increased GPU utilization from 25% → 75% through KV cache optimization and adaptive request coalescing, enabling near real-time inference on consumer hardware. • Reduced cost per 1,000 requests by 69% ($0.55 → $0.17) — translating hardware efficiency gains into direct production cost savings at scale. • Built three serving strategies (baseline, static batching, dynamic batching) using
  • M
    Multi-Agent Verified RAG System for Hallucination Reduction
    Nov 2025 - Jan 2026 (3 months)
    Built a multi-agent Retrieval-Augmented Generation (RAG) pipeline to reduce hallucinations in large language model outputs. • Implemented hybrid retrieval (BM25 + dense embeddings) with claim decomposition and NLI-based verification • Reduced hallucination rate from 50% (Naive RAG) to 0%, achieving 100% supported-claim precision across 58 evaluated claims • Developed a FastAPI evaluation backend to automate execution of Naive, Standard, and Verified RAG pipelines • Logged evaluation metrics and generated comparative performance charts for systematic benchmarking
  • E
    Emotion-Based Human–Computer Interaction (Facial Emotion Recognition)
    Sep 2025 - Nov 2025 (3 months)
    Built an end-to-end facial emotion recognition system enabling emotion-aware human–computer interaction using deep learning. • Implemented CNN, CBAM-attention CNN, and EfficientNet-B0 architectures for 8-class emotion classification • Achieved 66.49% test accuracy using EfficientNet-B0 with transfer learning — 5.2% improvement over baseline CNN models • Addressed class imbalance and performed comparative evaluation across multiple model architectures • Generated confusion matrices, ROC curves, and saliency maps to analyze model performance and interpretability • Explored real-world applications including adaptive user interfaces and mental wellness monitoring
This is a community-created genome.