Employment Details:
- Employment Type: Full-time.
- Work Mode: Hybrid (US-based, remote-friendly).
- Location: San Mateo, CA / New York, NY.
- Compensation: $176K - $224K Base (OTE: $220K - $280K).
- Seniority: 3+ years of experience.
Compensation and Benefits:
- Salary: $176K - $224K Base.
- OTE: $220K - $280K.
- Variable component paid quarterly based on individual and team performance.
- Compensation scales with experience.
- Candidates with 10+ years may be considered for above-range packages.
- Meaningful equity included on top of OTE.
Seniority Requirements:
- 3+ years of experience in customer-facing AI/ML field engineering (FDE, Applied AI, Solutions Architect, AI Infra, ML Engineer, Software Engineer with pre-sales exposure, or research backgrounds transitioning to customer-facing roles).
Work Experience:
- Shipped AI/ML production code inside a customer's environment.
- Hands-on LLM inference and fine-tuning experience (ran SFT pipelines, benchmarked latency, and tuned open-model deployments).
- Ran the full field cycle in a pre-sales or customer-facing capacity (discovery, POC scoping, load tests, evals, and model selection).
- Background at an AI-native/AI-infra startup (inference, MLOps, developer tooling) or enterprise SaaS with built-in AI features.
Hard Skills:
- LLM serving frameworks (vLLM, SGLang, TensorRT-LLM).
- Agents.
- Inference trade-offs.
- Terminal-comfortable.
- Python and Kubernetes proficiency.
- Trained open models and familiar with fine-tuning methodologies (SFT, DPO, RFT).
- GPU optimization for LLM workloads.
Soft Skills:
- Demonstrated executive presence in enterprise customer-facing roles.
- Navigated enterprise org politics end-to-end (champions, detractors, security reviews, and procurement cycles).
Miscellaneous:
- Domestic travel to enterprise customers as needed.
Key Requirements:
- Deep hands-on experience with LLM inference and/or training.
- Working knowledge of open-model frameworks (vLLM, SGLang, TensorRT-LLM) and fine-tuning workflows (SFT at minimum; DPO/RFT a strong plus).
- Proven ability to ship production code inside a customer's environment.
- Built and deployed POCs/MVPs that ran in someone else's production system.
- Strong Python skills plus GPU/cloud infrastructure experience (AWS, Azure, or GCP).
- Comfort with Kubernetes.
- Executive presence and enterprise navigation skills.
- Able to run a technical deep-dive with an ML engineer and present architecture trade-offs to a VP in the same afternoon.
- Pre-sales or customer-facing field engineering experience (FDE, Applied AI Engineer, Solutions Architect, or similar).
Tech Stack:
- Python.
- vLLM.
- SGLang.
- TensorRT-LLM.
- Kubernetes.
- AWS.
- Azure.
- GCP.
- Azure AI Foundry.
- AWS Bedrock.
- AWS SageMaker.
- GCP Vertex AI.
- LLM Fine-Tuning (SFT, DPO, RFT).
- GPU Infrastructure.
- Open-source LLM frameworks.