Model Evaluation and Threat Research, Inc. Logo

Model Evaluation and Threat Research, Inc.

Task Development Engineer

Posted 26 Days Ago
Remote
Hiring Remotely in USA
150-300 Hourly
Mid level
Remote
Hiring Remotely in USA
150-300 Hourly
Mid level
Develop novel, difficult evaluation tasks for frontier AI models; verify task specifications and solvability; baseline and score model and human completions; and improve task development infrastructure and workflows to support METR's Time Horizons evaluations.
The summary above was generated by AI
About METR

We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.

METR has consistently set precedents for catastrophic AI risk evaluations, including the first independent safety evaluations (working informally with Anthropic and OpenAI in 2022), the first loss-of-control evaluations and first agentic dangerous capability evaluations, the first evaluations using finetuning (mentioned briefly here), the first independent evaluations using internal information about training, the first review partnership for company risk analysis, the first embedded redteaming, and first evaluations of internal deployments, and the first independent misalignment incident investigation.

We’ve been consulted and/or favorably referenced by groups on opposite ends of various spectra, including a16z, Khosla, Gary Marcus, Obama, and Dean Ball, and are known for producing one of the most positive results on AI capabilities (the time horizon trend) and the most negative (our downlift study). We’re generally referenced as the canonical third party assessor, e.g. as the obvious candidate to verify conditional pause agreements, and are trusted with AI incident investigations by frontier labs and governments. 

We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.

About the role

  • As part of informing the world about risk from frontier AI systems, METR often runs and publishes evaluations of frontier models.
  • Time Horizons is a central tool the world uses to understand AI progress. Our methodology has been included in system cards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy. It is essential to our broader risk assessment work to have good capability evaluations.

  • Task Development Engineers contribute to METR’s expanding ambition of our evaluations with high quality tasks supporting the Time Horizons methodology. We expect our results to be seen by policymakers, frontier labs, national security stakeholders, and other key decisionmakers influencing society’s response to AI progress.

What this role looks like

  • (Primarily, and most importantly) Developing difficult, novel tasks for models. You will build well-scoped tasks that remain challenging as model time horizons grow, potentially to hundreds of hours.

  • Quality assurance for existing tasks. Once a task has been developed, you will verify that it's actually solvable as specified, and that the model is given (only) the information it needs.

  • Baselining and scoring tasks. Where helpful, you may be asked to baseline tasks within your domain of expertise, and/or score task completions from AIs or human baseliners.

  • Improving task development infrastructure. We're always improving our processes. Strong candidates will notice when existing workflows are inefficient or produce low-quality output, and take responsibility for improving them.

Skills we're looking for

  • Software engineering: You have several years of experience working on complex projects and codebases.

  • Evaluations: You have experience building hard (ideally agent-based) AI evaluations (e.g. RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA), ideally using the Inspect framework.

  • High attention to detail: You read closely, spot misspecifications and ambiguity, and pay attention to fiddly minutiae.

  • (Nice to have) Familiarity with METR infrastructure: Prior experience with Hawk, and familiarity with the methodology behind our Time Horizons work, is a plus.

Job details and compensation

  • Location: Remote (worldwide)
  • Hours: 20-40 hours per week (flexible schedule determined by you)
  • Timezone Requirements: A minimum of 1 hour (and ideally 4 hours) of overlap with the Pacific Coast Time workday, but you determine your exact work schedule.
  • Employment type: Contract / freelance
  • You decide the manner in which you complete your tasks to a standard that matches other professionals in this field.
  • Compensation: $150-300/hour. 
  • Top of this range is reserved for exceptional candidates.
  • Individuals who contribute >80 hours will be acknowledged in the final research output (if desired).

Our Culture
 
METR is a mission-driven organization. We believe our work can meaningfully shape humanity's future for the better, and we want to be the best people in the world doing this work. We have a tight-knit, collaborative research culture rooted in truth-seeking and integrity. We're fiercely committed to producing high-quality, trustworthy science. We're honest and transparent about our results, especially when they may go against the grain. We've earned trust as reliable partners who handle confidential information with care. We maintain a low-ego, drama-free environment focused on what matters.
 
We encourage you to apply even if your background may not seem like the perfect fit! We would rather review a larger pool of applications than risk missing out on a promising candidate for the position.
 
We are committed to diversity and equal opportunity in all aspects of our hiring process. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We welcome and encourage all qualified candidates to apply for our open positions.

Similar Jobs

An Hour Ago
Remote or Hybrid
95K-120K Annually
Senior level
95K-120K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Supports complex and non-standard sales transactions across the Americas by advising Sales on deal structure, pricing, commercial terms, approvals, quotes, revenue recognition, and quote-to-cash processes. Reviews deals for accuracy and policy compliance, escalates risks, and identifies process, system, and policy improvements. Partners with Sales, Finance, Legal, Revenue Accounting, Product Management, and Business Systems to improve deal velocity, pricing discipline, and data accuracy.
Top Skills: AICpq ToolsExcelSalesforceSalesforce Cpq
An Hour Ago
Remote or Hybrid
United States
95K-125K Annually
Expert/Leader
95K-125K Annually
Expert/Leader
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead end-to-end executive recruiting for AVP, VP, and C-suite roles. Develop search strategies, source and assess senior candidates, build leadership pipelines, and advise business leaders and HR stakeholders. Conduct market mapping, provide compensation and competitive insights, manage executive search firms, and deliver a high-touch candidate experience. Support onboarding, promote diverse candidate slates, and ensure recruiting practices align with company values and governance standards.
An Hour Ago
Remote or Hybrid
USA
85K-120K Annually
Entry level
85K-120K Annually
Entry level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Investigate and respond to real-time security detections across endpoint, identity, email, network, and cloud environments. Analysts perform forensic and malware analysis, assess Microsoft 365 incidents, identify root cause and compromise scope, recommend remediation, improve detection processes, and communicate findings clearly to customers. The role requires incident response, threat analysis, operating system, network, identity, SIEM, scripting, and AI fundamentals, with a four-day, ten-hour schedule including one weekend day.
Top Skills: Active DirectoryBashEdrLinuxmacOSMicrosoft 365Microsoft Entra IdMitre Att&CkOktaPowershellPythonSIEMWindows

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account