I cut small AWS charges by picturing them at 100x scale
In June 2026, I started reviewing my AWS bill. The goal was not to shave a few dollars off this month. Even a charge that is only a few dollars toda…
Tech news from the best sources
In June 2026, I started reviewing my AWS bill. The goal was not to shave a few dollars off this month. Even a charge that is only a few dollars toda…
The Silent Budget Killer Cloud bills creep up. You start with a small instance, a managed database, and a bucket. A year later, you're paying for re…
Backblaze B2 vs Self-Hosted S3: Which Actually Saves More Money in 2026? Backblaze B2's pricing is about as simple as cloud storage gets: $0.00695/G…
Most of us pick a model the same way: read a leaderboard, pick the best one we can afford, ship it. Then the bill arrives and the "cheap" model turn…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. Here's the pitch for best-of-N: instead of tr…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. The token bill is the cost you can see, becau…
Originally published on Loop & Retry — field notes on building LLM agents that survive production. Most fine-tuning guides answer "how many exam…
Bottom line: if you want one API key across OpenAI, Claude and Gemini and you're choosing mainly on token cost, put a thin router in front of your a…
Your shared Kubernetes cluster costs $80k/month. Which team owes what? If your answer is 'I don't know,' you have a finops problem that's about to b…
Your shared Kubernetes cluster costs $80k/month. Which team owes what? If your answer is 'I don't know,' you have a finops problem that's about to b…
xAI has a bold pitch for Grok 4.5: it codes about as well as Claude Opus 4.8, but does it with roughly a quarter of the tokens. That's not a "we're…
A 500 is an unexpected jolt from your sleep, while a 200 silently drains your savings. That's the bug you never get trained for. All is functioning.…
A shared AI API key feels fast when a team is still experimenting. One teammate builds a customer demo. Another wires an internal support bot. Someo…
Autoscaling is supposed to save money. A lot of the time it doesn't, because people reach for the wrong kind. Kubernetes gives you three common opti…
LLM features are cheap to prototype and surprisingly expensive to run at scale. A demo that costs pennies becomes a five-figure monthly bill once re…
How to Control CloudWatch Logs Costs on ECS? Originally published at https://fortem.dev/blog/cloudwatch-costs-ecs ECS sends all logs to CloudWatch w…
Tokenmaxxing Is a 2026 Anti-Pattern: Why Your Team's Token Bill Is Up 10x and What to Cut First There's a word floating through engineering Twitter…
Book: LLM Observability Pocket Guide: Picking the Right Tracing & Evals Tools for Your Team Also by me: Thinking in Go (2-book series) — Complet…
Book: LLM Observability Pocket Guide: Picking the Right Tracing & Evals Tools for Your Team Also by me: Thinking in Go (2-book series) — Complet…
The first number anyone quotes when asked what generative AI costs is a per-token figure. It is a comfortable number — small, unambiguous, available…
I get asked for receipts on every cost number I publish. So here is one full run, end to end, with screenshots replaced by file paths you can read o…
OpenSCAP with SOPS: The Hidden Cost of Supply Chain for Production Modern production environments rely heavily on automated compliance and secrets m…
The Hidden Cost of Scaling with Istio 1.20 and OpenShift: Benchmark Results Service mesh adoption has accelerated as organizations shift to microser…