Anthropic Fellows Program Fellow (AI Safety) at Anthropic | Torre

Anthropic Fellows Program Fellow (AI Safety)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation USD3.85k
location_on
Remote (for United States residents)
Remote (for United Kingdom residents)
Remote (for Canada residents)
Shared by
Emma of Torre.ai
1 day ago

Responsibilities


Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent - regardless of previous experience.Fellows will primarily use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). In one of our earlier cohorts, over 80% of fellows produced papers.What to expect4 months of full-time researchDirect mentorship from Anthropic researchersAccess to a shared workspace (in either Berkeley, California or London, UK)Connection to the broader AI safety and security research communityWeekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country)Funding for compute (~$15k/month) and other research expensesInterview processThe interview process will include an initial application & reference check, technical assessments & interviews, and a research discussion.We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed.Fellows workstreamsWe expect there to be significant overlap in the types of skills and responsibilities across the roles and will by default consider candidates for all the workstreams.You can see an overview of the current workstreams below:AI Safety FellowsAI Security FellowsML Systems & Performance FellowsReinforcement Learning FellowsEconomics & Societal Impacts FellowsAcross the workstreams, you may be a good fit if you:Are motivated by making sure AI is safe and beneficial for society as a wholeAre excited to transition into empirical AI research and would be interested in a full-time role at AnthropicHave a strong technical background in computer science, mathematics, or physicsThrive in fast-paced, collaborative environmentsCan implement ideas quickly and communicate clearlyStrong candidates may also have:Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)Experience in areas of research or engineering related to their workstreamCandidates must be:Fluent in Python programmingAvailable to work full-time on the Fellows programUnique candidate criteriaYou might be a particularly great fit for this workstream if you:Are motivated by reducing catastrophic risks from advanced AI systemsHave experience with empirical ML research projectsHave experience working with large language modelsHave experience in one of the research areas mentioned aboveHave a track record of open-source contributionsMentors, research areas, & past projectsFellows will undergo a project selection & mentor matching process. Potential mentors include:Sam BowmanSara PriceAlex TamkinNina PanicksseryTrenton BrickenLogan GrahamJascha Sohl-DicksteinJoe BentonCollin BurnsFabien RogerSamuel MarksKyle FishEthan PerezOur mentors will lead projects in select AI safety research areas, such as:Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.Adversarial Robustness and AI Control: Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios.Model Organisms: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.Model Internals / Mechanistic Interpretability: Advancing our understanding of the internal workings of large language models to enable more targeted interventions and safety measures.AI Welfare: Improving our understanding of potential AI welfare and developing related evaluations and mitigations.On our Alignment Science and Frontier Red Team blogs, you can read about past projects, including:Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data: Alex Cloud and Minh Le, et al., mentors including Samuel Marks and Owain EvansOpen-source circuits: Michael Hanna and Mateusz Piotrowski with mentorship from Emmanuel Ameisen and Jack LindseyFor a full list of representative projects for each area, please see these blog posts: Introducing the Anthropic Fellows Program for AI Safety Research, Recommendations for Technical AI Safety Research Directions.CompensationThe expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension).LogisticsTo participate in the Fellows program, you must have work authorization in the US, UK, or Canada and be located in that country during the program.We have designated shared workspaces in London and Berkeley where fellows will work from and mentors will visit. We are also open to remote fellows in the UK, US, or Canada. We will ask you about your availability to work from Berkeley or London (full- or part-time) during the program.We are not currently able to sponsor visas for fellows. To participate in the Fellows program, you need to have or independently obtain full-time work authorization in the UK, the US, or Canada.The program runs for 4 months, full-time. If you can't commit to the full duration, please still apply and note your constraints in the application. We review these requests on a case-by-case basis.Please note: We do not guarantee that we will make any full-time offers to fellows.