Frontier AI labs still won’t say how they’d contain a rogue model
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems…
Tech news from the best sources
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems…
Anthropic researchers found AI agents can clash, collude and coordinate in unexpected ways, raising new questions about whether today’s safety tests…
Vector icons do not line up with text out of the box. Unlike letters, graphic paths lack a consistent baseline, ascender, or descender. If you place…
What does a world of total user-aligned AI actually look like?
In Nicomachean Ethics VI.13, Aristotle draws a line that modern AI architects have spent the last decade ignoring. Epistēmē (ἐπιστήμη) is scientific…
RLHF vs DPO vs IPO vs KTO: which alignment method should you use You have a base model, say Llama 3.2 8B, that can write poetry in any meter and pas…
On fitting an AI with a listening hood. Prologue: This Is Not a Story About the Future When people talk about the risks of AI, one thought experimen…
Introduction For the last year and a half, I have been building SAFi (the Self-Alignment Framework Interface). It is a self-hosted, fully open-sourc…