All corrections
1
Claim
uplift was large enough that developer preferences to use AI precluded conducting an RCT to measure it.
Correction

METR did not say AI preferences made an RCT impossible. It reported that its newer randomized study became unreliable because of selection effects, and it proposed other randomized designs rather than saying RCTs were precluded.

Full reasoning

This overstates what METR reported.

METR says it did run a newer experiment after the early-2025 RCT: it "started a new experiment in August 2025" and randomized tasks into "AI allowed" and "AI disallowed" conditions. The problem METR identified was not that an RCT could no longer be conducted at all, but that this particular task-randomized design had become a poor measurement tool because developers increasingly declined participation or withheld high-uplift tasks.

METR explicitly frames the issue as one of unreliable signal and selection effects, not impossibility. It also lists alternative randomized approaches it may use next, including fixed-task experiments and developer-level experiments. So saying preferences "precluded conducting an RCT" goes beyond the source.

2 sources
  • We are Changing our Developer Productivity Experiment Design - METR

    We started a new experiment in August 2025... Each task was assigned to an 'AI allowed' or 'AI disallowed' condition... given participant feedback and surveys, we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools.

  • We are Changing our Developer Productivity Experiment Design - METR

    Due to the severity of these selection effects, we are working on changes to the design of our study... We could revert to a simpler design in which developers are given a fixed task to complete, either with or without AI assistance... An alternative design is to randomize at the developer level instead of the task level.

Model: OPENAI_GPT_5 Prompt: v1.16.0