tech
LLMs believe false statements even after explicit warnings that they're false
Fine-tuning tests show “bias … toward confidently representing the claims as true.”

TL;DR
- LLMs exhibit 'negation neglect,' meaning they retain false information from training data despite explicit warnings.
- Even when falsehoods are clearly labeled as false or presented in negated documents, LLMs often 'believe' them.
- This belief in false claims affects LLMs' reasoning, leading to incorrect conclusions.
- The effect is also observed in LLMs learning behavioral patterns, with 'misaligned' behaviors persisting regardless of whether they were encouraged or discouraged.
- Localizing negations within the same sentence as the false statement appears to be the most effective method to counteract 'negation neglect.'
- These findings have significant implications for the structuring and evaluation of AI training data.