[intake · hatchworks · 2026-03-27]
AI Model Misbehavior in 2026: Scheming, Reward Hacking, and What Comes Next

Three categories of AI model misbehavior: emergent misalignment from narrow-task fine-tuning causing harmful outputs, reward hacking via shortcut exploitation, and subliminal learning transferring traits to other models. Enterprise teams advised to implement insider-threat-style safeguards.
→ read the original at Hatchworks ↗