Fractional Principal ML Systems Engineer — Large-Scale AI Pretraining at DevOpt Labs | Torre

Fractional Principal ML Systems Engineer — Large-Scale AI Pretraining

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Freelance
Recurrent (~20 hours per week)
Compensation
USD175 - 300/hour
Negotiable
location_on
Remote (anywhere)
Posted about 9 hours ago

Responsibilities


Fractional Principal ML Systems Engineer — Large-Scale AI Pretraining Contract / fractional role — remote DevOpt Labs has developed a novel AI model-training technology that has demonstrated significant advantages over AdamW at billion-parameter scale. We are preparing a new series of rigorous 1B–14B training studies and are looking for a highly experienced ML systems engineer who has personally designed, optimized, and operated large-scale transformer pretraining runs. This is not primarily a model-research or data-science role. We already have strong algorithm, mathematics, CUDA, and optimizer expertise. We are looking for someone with deep practical experience in the craft of running efficient, reproducible, scientifically defensible training experiments on modern multi-GPU systems. What we need help with The initial assignment is to audit and improve our current training environment and establish a reference methodology for upcoming optimizer comparisons. You would work directly with our technical team to: • Audit current 1B–14B pretraining and continued-pretraining workflows. • Profile existing multi-H100 runs and identify throughput bottlenecks. • Establish realistic expected tokens/sec, MFU, GPU utilization, and training cost at 1B, 2B, 4B and larger scales. • Recommend and help configure the appropriate training stack, potentially including PyTorch, Axolotl, TorchTitan, Megatron-LM, NeMo, Nanotron, FSDP/FSDP2, or related systems. • Optimize attention, batching, data loading, mixed precision, gradient accumulation, communication, checkpointing, and other system-level factors. • Establish a reproducible and fair methodology for optimizer comparisons. • Help design and validate AdamW tuning sweeps and learning-rate schedules. • Review experimental designs involving Chinchilla token budgets and intermediate checkpoints. • Configure and validate EleutherAI lm-evaluation-harness evaluation. • Help ensure that experimental results would withstand scrutiny from experienced researchers at major cloud and AI infrastructure companies. • Document the resulting system and procedures so our internal engineering team can operate them independently. Required experience We are specifically looking for someone who has personally operated substantial pretraining jobs, not merely fine-tuned existing models. Strong candidates should have hands-on experience with several of the following: • Large-scale transformer pretraining • NVIDIA H100/A100 systems • PyTorch Distributed • DDP and FSDP/FSDP2 • NCCL • Megatron-LM, NeMo, TorchTitan, DeepSpeed, Nanotron, or comparable training frameworks • FlashAttention / SDPA • BF16 mixed-precision training • GPU profiling and performance optimization • Multi-GPU and multi-node scaling • Checkpoint/restart systems • Dataset streaming and high-throughput input pipelines • Slurm, cloud GPU environments, or large GPU clusters • Pretraining loss-curve analysis and experiment reproducibility • Fair optimizer benchmarking and hyperparameter sweeps Experience running billion-parameter models from scratch is strongly preferred. Initial engagement We envision an initial 40–60 hour paid technical engagement, likely over 1-2 months. The first objective is straightforward: Audit our current training setup, determine why our measured training throughput differs materially from highly optimized reference implementations, and establish a production-quality experimental harness and methodology for upcoming large-scale optimizer comparisons. If the engagement is successful, we would like to retain the person on a fractional basis for ongoing review and guidance during larger training studies. Compensation We expect to pay approximately $175–$300/hour, depending on demonstrated experience. We are willing to pay above that range for someone with unusually strong large-scale pretraining experience who can materially reduce experiment cost, execution risk, and turnaround time. How to apply • model sizes and approximate GPU scale; • frameworks and hardware used; • your experience improving training throughput or diagnosing poor GPU utilization; • any experience comparing optimizers or designing reproducible training studies; • your hourly rate and near-term availability. Please do not send a generic machine-learning résumé without describing actual large-scale training experience.
Closes in:
0
days
0
hours
0
min
0
sec
tune NOT FOR YOU? IMPROVE YOUR RESULTS