Machine Learning Ops Engineer at Stratum Ai | Torre

Machine Learning Ops Engineer

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Provide your expected compensation while applying
location_on
Remote (for Canada residents)
Shared by
Emma of Torre.ai
12 days ago

Responsibilities


We are looking for a high-agency Machine Learning Ops Engineer to join our Infrastructure Team. You will help build and maintain the platform used to train, evaluate, and serve our AI models to clients in the mining industry. Your work will directly support our Technical Services and Platform teams in delivering solutions that create value for mining clients.This position requires strong expertise in Python and machine learning workflows. You will work alongside a team of three engineers focused on creating robust infrastructure and tooling.This is a remote-first position based in Canada.Key ResponsibilitiesDevelop robust and well-tested code for core internal tools:Create data preprocessing modules for mining dataImplement metrics calculations and evaluation pipelinesBuild visualization tools for 3D models and ML performance metricsTroubleshoot and fix issues in existing metrics codeBuild and maintain our custom end-to-end MLOps platform:Implement experiment tracking systemsCreate model registry with versioning and storageDevelop automated testing frameworksBuild interfaces between different components of the ML pipelineDevelop production-grade QA/QC systems for deployed AI models:Implement input data validationCreate automated alerts for performance issuesSet up monitoring for data driftBuild dashboards for model performance metricsCreate specialized tools for mining data:Implement spatial data processing utilitiesBuild visualization tools for 3D geological dataDevelop data converters between different mining data formatsCreate utilities for coordinate transformationsRefactor and productionize code created by the client services team:Convert notebooks into modular Python packagesImplement proper error handling and loggingAdd comprehensive testing to existing codeImprove performance of data processing pipelinesProvide technical expertise to the client services teamManage infrastructure for data processing, model training, and servingMentor junior engineers, perform code reviews, and write documentationProactively identify technical challenges and drive improvement initiativesTechnical Competencies & RequirementsBachelor's degree in Computer Science, Engineering, or related fields OR equivalent experience in software development and ML engineering3+ years of industry experienceKubernetes, PyTorchAdvanced Python programming skills:Proficiency with data science libraries (numpy, pandas)Experience with visualization toolsAbility to write modular, robust, and tested Python codeStrong debugging skills for complex ML systemsDeep learning experience:Implementation of neural network models and training workflowsUnderstanding of model architecture selectionKnowledge of model evaluation techniquesMLOps expertise:Creating experiment tracking systemsBuilding model registries and versioning systemsImplementing model deployment pipelinesSetting up monitoring for model performanceData engineering capabilities:Experience with SQL and database principlesFamiliarity with database frameworksAbility to create data processing pipelinesExperience handling common mining data formats and transformationsInfrastructure management:Experience with cloud services (AWS/Azure)Understanding of containerization (Docker or Singularity)Knowledge of compute resources for MLTesting and quality assurance:Implementing automated tests for ML systemsCreating QA/QC systems for model predictionsDesigning validation steps for data inputs/outputsAbility to write efficient software following best practicesProven ability to thrive in startup environments with low structure and high autonomyStrong technical communication skills and ability to collaborate in a remote team settingExperience working with machine learning in computer vision, NLP, recommender systems, or scientific applicationsStrong background in probability, machine learning, and data scienceStrong experience with data analysis/processing libraries such as pandas and numpyExcellent communication skills for both technical and non-technical audiencesSelf-learner and motivated to pick up new skillsNice to HavePrevious experience working at startupsFamiliarity with Git, experiment tracking tools (WandB, Comet, etc.)Experience working on production machine learning using tools such as KubeFlow, MLFlow, AirFlow, Seldon Core, DVC, Spark, etc.Written/oral fluency in a language besides EnglishExperience optimizing data processing pipelines and/or neural network modelsProficiency in a lower-level programming language or GPU programmingExperience with data application frameworksFull stack development experienceExperience with experiment tracking systems and ML model monitoringBackground in mining or resource modelingAbout StratumWe're Stratum, a mining software company with machine learning models as our core product. Our 3D maps predict how gold, silver, copper, etc. are distributed (and how much!) using only small amounts of data, unconventional data processing, and proprietary ML protocols. Our work directly affects how much money a mine is going to make next week/month/year while reducing waste/cost. We're supported by Founders Fund, Aramco, Builders VC, Y Combinator, and Ilya Sutskever, former Chief Scientist at OpenAI, who have recognized the potential of our industry-disrupting technology.Our long-term vision is to build a massive AI engine capable of making every decision in a mining operation, down to moving individual rocks. If you’re an exceptional engineer interested to helping make this vision a reality we look forward to reviewing your application and working together.