Scientists from a British artificial intelligence cybersecurity firm called Mindgard claim to have developed a trick that allows them to convince ChatGPT to create dangerous images with just one command: ChatGPT unsafe image generation. In their words, only minor alterations to the already popular internet prompt used for generating funny images were made. However, after that, ChatGPT began creating controversial material.
Researchers Find ChatGPT Can Be Made to Generate Inappropriate Images
The researchers shared their findings with the BBC. They said the latest public version of ChatGPT could generate images that included graphic and inappropriate content even when the prompt did not clearly ask for it. Because of this, they believe more work is needed to improve AI safety systems.
Read more: OpenAI Introduces “Trusted Contact” Feature in ChatGPT for User Safety Support
OpenAI Responds to the Findings
After being contacted by the BBC, OpenAI said it investigated the issue and added new protections. The company said it introduced extra safeguards to stop users from getting those types of images through similar prompts. OpenAI also explained that ChatGPT has several layers of protection designed to block content that breaks its rules. The company said it uses automated systems and human reviewers to help detect and stop harmful material. It added that it continues to improve these protections over time.
Researchers Say Problems Still Exist
Mindgard said that even after the new protections were added, small changes to the prompt could still produce concerning results. The researchers did not publicly reveal the exact prompt they used. However, they showed examples of the output to journalists. Peter Garraghan, the founder of Mindgard and a professor at Lancaster University, said the most worrying part was that the prompt looked harmless. He explained that the instruction did not ask for specific harmful topics, yet the AI still created disturbing images on its own. He said this showed that some weaknesses remain in current AI systems, like ChatGPT’s 1 billion monthly active users.
Impact on Researchers
The company said one of its researchers was deeply affected by the images generated during testing. AI safety researcher Jim Nightingale said some of the content was very disturbing. According to Mindgard, the images included scenes involving injuries, crime-related situations, and other graphic material. The work of Mindgard emphasizes “red teaming,” which is a methodology that uses experts to identify vulnerabilities in AI systems. The idea is to assist technology firms in identifying their flaws and solving them before they impact more users, such as the ChatGPT banking finance feature.
Read more: Yango Ride Introduces ChatGPT Integration for In-Chat Trip Planning in Over 25 Countries
Why AI Safety Remains a Challenge
Experts say preventing all harmful AI outputs is difficult. Large language models are trained on huge amounts of data collected from the internet. Because of this, AI systems can sometimes respond in unexpected ways. Dr. Rumman Chowdhury, an AI evaluation expert who was not involved in the research, said improving AI safety is a constant challenge. She explained that AI models do not truly understand meaning, intent, or right and wrong the way humans do. As companies improve protections, new methods are often found to get around them.
It is clear that the problem of ensuring safety of AI systems remains open. According to OpenAI, some additional safety features have already been added and monitoring continues. Scientists appreciate the company’s efforts in the area but are sure that additional changes should be made. The more efficient the AI systems become, the harder it becomes to make them safe enough. News source eTimes Pakistan.

