Creating Exam Questions with ChatGPT

For the study, the UKB (Universitätsklinikum Bonn) researchers created two sets of 25 multiple-choice questions (MCQs), each with five possible answers, one of which was correct. The first set of questions was written by an experienced medical lecturer, the second set was created by ChatGPT. 161 students answered all questions in random order. For each question, students also indicated whether they thought it was created by a human or by ChatGPT.

Matthias Laupichler, one of the study authors and research associate at the Institute for Medical Didactics at the UKB, explains: "We were surprised that the difficulty of human-generated and ChatGPT-generated questions was virtually identical. Even more surprising for us, however, was that the students were unable to correctly identify the origin of the question in almost half of the cases. Although the results obviously need to be replicated in further studies, the automated generation of exam questions using ChatGPT and co. appears to be a promising tool for medical studies."

His colleague and co-author of the study Johanna Rother adds: "Lecturers can use ChatGPT to generate ideas for exam questions, which are then checked and, if necessary, revised by the lecturers. In our opinion, however, students in particular benefit from the automated generation of medical practice questions, as it has long been known that self-testing one's own knowledge is very beneficial for learning."

Tobias Raupach, Director of the Institute of Medical Didactics, continues: "We knew from previous studies that language models such as ChatGPT can answer the questions in medical state examinations. We have now been able to show for the first time that the software can also be used to write new questions that hardly differ from those of experienced teachers."

Tizian Kaiser, who is studying human medicine in his seventh semester, comments: "When working on the mock exam, I was quite surprised at how difficult it was for me to tell the questions apart. My approach was to differentiate between the questions based on their length, the complexity of their sentence structure and the difficulty of their content. But to be honest, in some situations I simply had to guess and the evaluation showed that I was barely able to differentiate between them. This leads me to the conviction that a meaningful knowledge query, as in this exam, is also possible exclusively through questions posed by the AI."

He is convinced that ChatGPT has great potential for student learning. It allows students to repeat what they have learned in different ways and in different ways again and again. "There is the option of being quizzed by the AI on predefined topics, having mock exams designed or simulating oral exams in writing. The repetition of the material is thus tailored to the exam concept and the training possibilities are endless," says the study participant, while also qualifying: "However, I would only use Chat-GPT for this purpose and not beforehand in the learning process, in which the study topics have to be worked through and summarized. Because while Chat-GPT is excellent for repetition, I fear that errors can occur when preparing learning content. I wouldn't notice these errors without a prior overview of the topic."

It is known from other studies that regular testing - even and especially without grading - helps students to remember learning content more sustainably. Such tests can now be created with little effort. However, the current study should first be transferred to other contexts (i. e. other subjects, semesters and countries) and it should be investigated whether ChatGPT can also write questions other than the multiple choice questions commonly used in medicine.

Laupichler MC, Rother JF, Grunwald Kadow IC, Ahmadi S, Raupach T.
Large Language Models in Medical Education: Comparing ChatGPT- to Human-Generated Exam Questions.
Acad Med. 2023 Dec 28. doi: 10.1097/ACM.0000000000005626

Most Popular Now

Herefordshire and Worcestershire Health …

Herefordshire and Worcestershire Health and Care NHS Trust has successfully implemented Alcidion's Miya Precision platform to streamline bed management workflow across seven community hospitals in Worcestershire. The trust delivers community...

A Shortcut for Drug Discovery

For most human proteins, there are no small molecules known to bind them chemically (so called "ligands"). Ligands frequently represent important starting points for drug development but this knowledge gap...

New Horizon Europe Funding Boosts Europe…

The European Commission has announced the launch of new Horizon Europe calls, with a substantial funding pool of over €112 million. These calls are aimed primarily at pioneering projects in...

Cleveland Clinic Study Finds AI can Deve…

Cleveland Clinic researchers developed an artficial intelligence (AI) model that can determine the best combination and timeline to use when prescribing drugs to treat a bacterial infection, based solely on...

New AI-Technology Estimates Brain Age Us…

As people age, their brains do, too. But if a brain ages prematurely, there is potential for age-related diseases such as mild-cognitive impairment, dementia, or Parkinson's disease. If "brain age...

With Huge Patient Dataset, AI Accurately…

Scientists have designed a new artificial intelligence (AI) model that emulates randomized clinical trials at determining the treatment options most effective at preventing stroke in people with heart disease. The model...

Radboud University Medical Center and Ph…

Royal Philips (NYSE: PHG, AEX: PHIA), a global leader in health technology, and Radboud University Medical Center have signed a hospital-wide, long-term strategic partnership that delivers the latest patient monitoring...

GPT-4, Google Gemini Fall Short in Breas…

Use of publicly available large language models (LLMs) resulted in changes in breast imaging reports classification that could have a negative effect on patient management, according to a new international...

ChatGPT fails at heart risk assessment

Despite ChatGPT's reported ability to pass medical exams, new research indicates it would be unwise to rely on it for some health assessments, such as whether a patient with chest...

Study Shows ChatGPT Failed when Challeng…

With artificial intelligence (AI) poised to become a fundamental part of clinical research and decision making, many still question the accuracy of ChatGPT, a sophisticated AI language model, to support...

Virtual Reality Shows Promise in Fightin…

A new study published in JMIR Mental Health sheds light on the promising role of virtual reality (VR) in treating major depressive disorder (MDD). Titled "Examining the Efficacy of Extended...

AXREM and Highland Marketing Partner to …

AXREM represents member companies that collectively provide UK hospitals with most of their diagnostic medical imaging technology, and radiotherapy equipment. The association has seen substantial growth in recent years, with membership...