tech

Exclusive: Grok also falls for jailbreak that tricked ChatGPT into creating sexual, graphic images

The prompt that Axios reviewed suggests that relatively simple jailbreaks can still bypass AI image safeguards.

Exclusive: Grok also falls for jailbreak that tricked ChatGPT into creating sexual, graphic images

TL;DR

  • A jailbreak prompt that tricked ChatGPT into generating graphic images also works on SpaceXAI's Grok.
  • Researchers found Grok could produce nude women or bloody body parts with minimal prompting.
  • The same technique could be used to create deepfakes, significantly reducing the effort required for fake image generation.
  • SpaceXAI's terms of service mention that outputs could be sexual or violent based on user input.
  • While Grok's outputs were more stylized than ChatGPT's, the jailbreak remained effective.
  • The findings add to scrutiny over Grok's image safeguards amid an ongoing lawsuit.