About the RoleWe are looking for an LLM Engineer with strong hands-on expertise in LLM inference, fine-tuning, model evaluation, and production deployment.The role combines deep technical ownership with client interaction. You will be responsible for understanding client requirements, designing the right LLM solution, and taking projects from POC to production.Key ResponsibilitiesLLM EngineeringDesign and develop LLM solutions for real-world business use cases.Fine-tune LLMs using techniques such as SFT, LoRA, QLoRA, and PEFT.Prepare and validate datasets for fine-tuning and model evaluation.Configure and optimize LLM inference for quality, latency, throughput, scalability, and cost.Work with inference frameworks such as vLLM, TGI, Ollama, or equivalent.Optimize GPU utilization, memory usage, batching, quantization, and other inference settings.Evaluate model performance and continuously improve response quality, accuracy, and reliability.Integrate LLMs with RAG, APIs, databases, applications, and AI workflows.Client & Project DeliveryUnderstand client business and technical requirements and translate them into practical AI solutions.Lead technical discussions, solution walkthroughs, demos, and POCs with clients.Own the end-to-end technical delivery of assigned LLM projects.Coordinate with Product and Engineering teams to deliver projects within agreed scope and timelines.Identify technical risks, dependencies, and performance issues and drive them to resolution.Deploy, monitor, and support LLM solutions in production.Required SkillsStrong hands-on experience with LLMs and Generative AI.Excellent understanding of LLM inference and inference optimization.Practical experience with LLM fine-tuning and model evaluation.Strong Python programming skills.Experience with PyTorch and Hugging Face Transformers.Experience with LoRA/QLoRA/PEFT or similar fine-tuning approaches.Understanding of GPUs, CUDA, memory management, and inference performance.Experience with LLM serving frameworks such as vLLM, TGI, Ollama, or equivalent.Understanding of RAG, embeddings, vector databases, and prompt engineering.Experience deploying AI solutions in production.Strong problem-solving and client communication skills.Good to HaveExperience with LLM quantization and model compression.Experience with multi-GPU or distributed inference.Experience with AWS, GCP, or Azure.Experience building AI agents and tool-calling systems.Experience working on enterprise/client AI projects.Share your CV at ceo@inbotiq.com