The Slow Lane: Latency Engineering When Your AI Endpoint Is Free
Free model access solves the cost problem and creates a latency problem, and most teams measure the wrong number. My position is direct: the p95 of…
Tech news from the best sources
Free model access solves the cost problem and creates a latency problem, and most teams measure the wrong number. My position is direct: the p95 of…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. The demo felt instant. The agent answered in…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. Here's the pitch for best-of-N: instead of tr…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. The token bill is the cost you can see, becau…
Recently, I've added a bunch of hype-monsters to my AI Werewolf : Kimi K3 Qwen 3.8 Max, Qwen 3.7 Plus, Qwen 3.7 Flash MiniMax M3 Plus the ones I've…