Presentation: Can Claude Fix Itself? Using LLMs for Incident Response
Anthropic reliability engineer Alex Palcuie shares practical lessons on using LLMs for real-world incident response. He explains where AI acts as a…
Tech news from the best sources
Anthropic reliability engineer Alex Palcuie shares practical lessons on using LLMs for real-world incident response. He explains where AI acts as a…
Cloudflare has recently detailed how it is using AI to transform internal engineering standards from passive documentation into an actively enforced…
The engineering team at Stripe recently described how they automated database incident recovery by modeling their global infrastructure as a graph.…
Instacart introduced Blueberry, an AI-assisted incident response system that helps on-call engineers investigate production issues faster. It combin…
Expedia Group has introduced STAR, an internal AI-assisted observability platform that helps engineers investigate production incidents using servic…
A configuration change in AWS's bill computation system showed customers estimated bills in the billions and trillions of dollars for over 24 hours.…
OpenAI found two unrelated bugs masquerading as one in ChatGPT's data infrastructure. Silent hardware corruption on one Azure host and an 18-year-ol…
To provide SRE as a service, a team built a center of excellence, introducing Federated SREs and roles like production manager and technical tribe l…
Google Cloud's automated systems suspended Railway's production account without notice, triggering an eight-hour platform-wide outage affecting 3 mi…
SRE часто внедряют как набор инструментов, дашбордов и новых должностей, но через полгода команда всё так же тушит инциденты по ночам, а бюджеты оши…
Platform engineering succeeds when reliability and ergonomics reinforce each other rather than compete. This article explores three foundational pil…