> All of this went down far too fast for a human-driven response. The only way… more
~ai.alignment×
10 links
> Kudos to OpenAI for sharing their recent experiences with a misaligned… more
> This incident occurred during an internal evaluation which prompts models to… more
> We use agentic misalignment as a case study to highlight some of the… more
> We introduce Natural Language Autoencoders (NLAs), an unsupervised method for… more
> As we wrote in the Project Glasswing announcement, we do not plan to make… more
> Claude hadn’t yet discovered it was in BrowseComp, but it had correctly… more
> A deeper look at confessions, reward hacking, and monitoring in alignment… more
> After recovering, Tan joined online support groups for other survivors of AI… more
> LLMs are useful because they generalize so well. But can you have too much of… more