SciCode Trainer (Material Science) - Scientific Coding Tasks in Python at turing | Torre

SciCode Trainer (Material Science) - Scientific Coding Tasks in Python

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Freelance
Recurrent (~40 hours per week)
Provide your expected compensation while applying
location_on
Remote (for Bangladesh residents)
Remote (for Brazil residents)
Remote (for Colombia residents)
Remote (for Egypt residents)
Shared by
Diana Montoya
5 days ago

Responsibilities


Role Overview: Turing is building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines.Responsibilities:Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solutionImplement verified golden solutions in Python with complete unit test coverageDesign discriminative test cases that clearly differentiate correct from incorrect model outputsRun QC validation checks on the Turing Central Task Platform (CTP) including Tier 1 structure checks and Tier 2 quality rubricsIterate on tasks based on QC feedback to meet Pass@K evaluation criteria across multiple LLM judges (GPT, Gemini, Nemotron)Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submissionParticipate in sync calls for reviews, feedback sessions, and project standups during overlap hoursRequired Qualifications:Master's or PhD in Material science or related field. Strong Python programming skills with experience in scientific computingAbility to write rigorous, well-posed scientific problems with clear constraints and expected outputsAttention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinismPrior experience in AI data annotation, research, or scientific writingFamiliarity with LLM evaluation frameworks or coding benchmarksExperience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific toolsPublished research or academic project experience in a STEM domainOffer Details:Commitments Required : Overlap of 4 hours with PST and 40 hrs/weekEngagement type : Contractor assignment/freelancer (no medical/paid leave)Duration of contract : 8 weeks