seantrott.substack.com/p/language-models-and-language-change-721
2 corrections found
The five that frequently used LLMs achieved an average accuracy of 92.7%;
The paper gives 92.7% as the experts’ true positive rate, not their overall accuracy.
Full reasoning
The cited paper does not report 92.7% as expert annotators’ overall accuracy. It reports TPR = 92.7% for experts, alongside a separate FPR value. In Table 1, the expert row is Avg. TPR 92.7 and Avg. FPR 4.0; the text below says the experts were "able to detect AI-generated text very reliably, achieving a TPR of 92.7%." So the article is mislabeling a detection-rate metric as accuracy.
1 source
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
Metric Nonexperts Experts Avg. TPR 56.7 92.7 Avg. FPR 51.7 4.0 ... the five annotators who have significant experience with using LLMs for writing-related tasks are able to detect AI-generated text very reliably, achieving a TPR of 92.7%.
the majority vote of these five achieved almost perfect performance (99.9%),
The paper reports the expert majority vote missed 1 of 300 articles overall, i.e. about 99.7% accuracy or 99.3% TPR overall—not 99.9%.
Full reasoning
The cited study says the majority vote of the five expert annotators misclassified only 1 of 300 articles. That is about 99.7% overall accuracy, not 99.9%. In the paper’s results table, the expert majority vote is also listed as 99.3% TPR with 0% FPR overall. So the article’s parenthetical figure of 99.9% does not match the paper’s reported overall results.
2 sources
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
In fact, the majority vote among five such "expert" annotators misclassifies only 1 of 300 articles...
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
Expert Majority Vote ... OVERALL TPR% (FPR%) ... 99.3 (0)