Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers

Сравнение научных аннотаций, сгенерированных ChatGPT, с оригинальными аннотациями с использованием детектора AI-выхода, детектора плагиата и слепых рецензентов
Alexander T. Pearson, Yuan Luo, Emma Dyer, Nikolay S. Markov, Catherine A. Gao, Frederick M. Howard, Siddhi Ramesh
2022-12-27

ChatGPTartificial intelligence output detectorblinded human reviewersplagiarism detectorscientific abstracts
Abstract Background Large language models such as ChatGPT can produce increasingly realistic text, with unknown information on the accuracy and integrity of using these models in scientific writing. Methods We gathered ten research abstracts from five high impact factor medical journals (n=50) and asked ChatGPT to generate research abstracts based on their titles and journals. We evaluated the abstracts using an artificial intelligence (AI) output detector, plagiarism detector, and had blinded human reviewers try to distinguish whether abstracts were original or generated. Results All ChatGPT-generated abstracts were written clearly but only 8% correctly followed the specific journal’s formatting requirements. Most generated abstracts were detected using the AI output detector, with scores (higher meaning more likely to be generated) of median [interquartile range] of 99.98% [12.73, 99.98] compared with very low probability of AI-generated output in the original abstracts of 0.02% [0.02, 0.09]. The AUROC of the AI output detector was 0.94. Generated abstracts scored very high on originality using the plagiarism detector (100% [100, 100] originality). Generated abstracts had a similar patient cohort size as original abstracts, though the exact numbers were fabricated. When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, but that the generated abstracts were vaguer and had a formulaic feel to the writing. Conclusion ChatGPT writes believable scientific abstracts, though with completely generated data. These are original without any plagiarism detected but are often identifiable using an AI output detector and skeptical human reviewers. Abstract evaluation for journals and medical conferences must adapt policy and practice to maintain rigorous scientific standards; we suggest inclusion of AI output detectors in the editorial process and clear disclosure if these technologies are used. The boundaries of ethical and acceptable use of large language models to help scientific writing remain to be determined.
1
AI output detector assigned high likelihood scores to generated abstracts: median 99.98% (IQR 12.73, 99.98) versus 0.02% (IQR 0.02, 0.09) for originals.
2
Blinded human reviewers correctly identified 68% of generated abstracts but mistakenly labeled 14% of original abstracts as generated; reviewers described generated abstracts as vaguer and formulaic.
3
ChatGPT generated clear scientific abstracts for 50 papers but only 8% adhered to specific journal formatting requirements.
4
Generated abstracts fabricated exact numerical data (e.g., patient cohort sizes) but matched the original abstracts' cohort sizes in magnitude.
5
Plagiarism detector found generated abstracts to be highly original, with 100% [100, 100] originality scores.
6
Recommendation: include AI output detectors in editorial processes and require clear disclosure when large language models are used for scientific writing.
7
The AI output detector achieved AUROC of 0.94 for distinguishing generated from original abstracts.

Scientific research abstracts (original vs ChatGPT-generated)

Comparative detection and characteristics of ChatGPT-generated versus original scientific abstracts using an AI output detector, a plagiarism detector, and blinded human reviewers, including formatting compliance, originality scores, fabricated data presence, and reviewer identification accuracy

Publication Details
Publication Date
2022-12-27
Journal
Publisher
ISSN
Cited by
419
Access Type
Author Information
Authors
Alexander T. Pearson
Yuan Luo
Emma Dyer
Nikolay S. Markov
Catherine A. Gao
Frederick M. Howard
Siddhi Ramesh
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%