Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers
Сравнение научных аннотаций, сгенерированных ChatGPT, с оригинальными аннотациями с использованием детектора AI-выхода, детектора плагиата и слепых рецензентов
2022-12-27
SCID: 54.1/sufwfm93
Discuss with AI
ChatGPTartificial intelligence output detectorblinded human reviewersplagiarism detectorscientific abstracts
Figures from the paper
Abstract (AI)
Abstract Background Large language models such as ChatGPT can produce increasingly realistic text, with unknown information on the accuracy and integrity of using these models in scientific writing. Methods We gathered ten research abstracts from five high impact factor medical journals (n=50) and asked ChatGPT to generate research abstracts based on their titles and journals. We evaluated the abstracts using an artificial intelligence (AI) output detector, plagiarism detector, and had blinded human reviewers try to distinguish whether abstracts were original or generated. Results All ChatGPT-generated abstracts were written clearly but only 8% correctly followed the specific journal’s formatting requirements. Most generated abstracts were detected using the AI output detector, with scores (higher meaning more likely to be generated) of median [interquartile range] of 99.98% [12.73, 99.98] compared with very low probability of AI-generated output in the original abstracts of 0.02% [0.02, 0.09]. The AUROC of the AI output detector was 0.94. Generated abstracts scored very high on originality using the plagiarism detector (100% [100, 100] originality). Generated abstracts had a similar patient cohort size as original abstracts, though the exact numbers were fabricated. When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, but that the generated abstracts were vaguer and had a formulaic feel to the writing. Conclusion ChatGPT writes believable scientific abstracts, though with completely generated data. These are original without any plagiarism detected but are often identifiable using an AI output detector and skeptical human reviewers. Abstract evaluation for journals and medical conferences must adapt policy and practice to maintain rigorous scientific standards; we suggest inclusion of AI output detectors in the editorial process and clear disclosure if these technologies are used. The boundaries of ethical and acceptable use of large language models to help scientific writing remain to be determined.
Key Findings
1
AI output detector assigned high likelihood scores to generated abstracts: median 99.98% (IQR 12.73, 99.98) versus 0.02% (IQR 0.02, 0.09) for originals.
2
Blinded human reviewers correctly identified 68% of generated abstracts but mistakenly labeled 14% of original abstracts as generated; reviewers described generated abstracts as vaguer and formulaic.
3
ChatGPT generated clear scientific abstracts for 50 papers but only 8% adhered to specific journal formatting requirements.
4
Generated abstracts fabricated exact numerical data (e.g., patient cohort sizes) but matched the original abstracts' cohort sizes in magnitude.
5
Plagiarism detector found generated abstracts to be highly original, with 100% [100, 100] originality scores.
6
Recommendation: include AI output detectors in editorial processes and require clear disclosure when large language models are used for scientific writing.
7
The AI output detector achieved AUROC of 0.94 for distinguishing generated from original abstracts.
Research Object
Scientific research abstracts (original vs ChatGPT-generated)
Research Subject
Comparative detection and characteristics of ChatGPT-generated versus original scientific abstracts using an AI output detector, a plagiarism detector, and blinded human reviewers, including formatting compliance, originality scores, fabricated data presence, and reviewer identification accuracy
Publication Details
Publication Date
2022-12-27
Journal
Publisher
ISSN
Cited by
419
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by15
Opinion Paper: “So what if ChatGPT wrote it?” Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy2023
ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns2023
Chatting and cheating: Ensuring academic integrity in the era of ChatGPT2023
ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?2023
Artificial Hallucinations in ChatGPT: Implications in Scientific Writing2023
The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic review2024
A SWOT analysis of ChatGPT: Implications for educational practice and research2023
A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly2024
ChatGPT for Education and Research: Opportunities, Threats, and Strategies2023
The role of ChatGPT in higher education: Benefits, challenges, and future research directions2023
Chatting and Cheating. Ensuring academic integrity in the era of ChatGPT2023
War of the chatbots: Bard, Bing Chat, ChatGPT, Ernie and beyond. The new AI gold rush and its impact on higher education2023
AI-generated research paper fabrication and plagiarism in the scientific community2023
Generative AI Literacy: Twelve Defining Competencies2024
Using ChatGPT for human–computer interaction research: a primer2023