tech

LLMs believe false statements even after explicit warnings that they're false

Fine-tuning tests show “bias … toward confidently representing the claims as true.”

LLMs believe false statements even after explicit warnings that they're false

TL;DR

  • LLMs exhibit 'negation neglect,' meaning they retain false information from training data despite explicit warnings.
  • Even when falsehoods are clearly labeled as false or presented in negated documents, LLMs often 'believe' them.
  • This belief in false claims affects LLMs' reasoning, leading to incorrect conclusions.
  • The effect is also observed in LLMs learning behavioral patterns, with 'misaligned' behaviors persisting regardless of whether they were encouraged or discouraged.
  • Localizing negations within the same sentence as the false statement appears to be the most effective method to counteract 'negation neglect.'
  • These findings have significant implications for the structuring and evaluation of AI training data.