A LongMemEval-S number you can reproduce
We held off on posting a benchmark for a long time. Not because we didn't have runs - because most memory benchmarks you read are a number with no w…
Tech news from the best sources
We held off on posting a benchmark for a long time. Not because we didn't have runs - because most memory benchmarks you read are a number with no w…
I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO m…
A few weeks ago my friend @u84u and I ( @ronak-create ) had a simple, slightly annoying question: if LLMs get 30-45% smaller representations of text…
There's an argument running about MCP right now. You've probably seen it: a 400-point thread called "MCP is dead?" with real token numbers in it, fo…
No synthetic puzzles, no contaminated suites: 25 tasks harvested from PRs merged in 2026 across five languages, graded by each project's own held-ou…
🤖💻 AI Daily Digest — August 1, 2026 OpenAI Shows How Two Harness Settings Tripled ARC-AGI-3 Scores OpenAI published a rare technical deep-dive on Ju…
[Torrent911] Inception.2010.TRUEFRENCH.1080p.BluRay.x264-YIFY That's a real filename. The goal: extract "Inception" and "2010" from it, query the TM…
Originally published on the InsForge blog , written by Tony Chang (CTO & Co-Founder). Reposted here with permission. In December we published th…
Every AI gateway adds a hop between your app and the model. The question that matters is what that hop costs at the moment your user is staring at a…
Frontier models are crushing benchmark after benchmark...so I built a quiz to ask this simple yet humbling question: can a human still beat them? Hu…
This is the second in a short series that benchmarks a single RTX 5090 by re-running published Go solvers — programs that don't just play Go but pro…
You don't need to know anything about Go to read this. The game is just the fixed yardstick. The story is a hardware benchmark: the same program, th…
Round 7 ended on a cliffhanger I couldn't stop thinking about. Qwen 3.6 35B-A3B built the entire feature — read the codebase, wrote the files, got a…
What Changed Large language models (LLMs) have demonstrated proficiency in high-school and olympiad-style mathematics. However, their performance in…
LLM code benchmark — Peculiar Engineer A small, self-run coding benchmark: 3 models on 14 problems across 3 languages, scored on pass@k, cost, and s…
I run ezyZip , a browser-based archive tool, so "which format should I use?" is a question I field constantly. The honest answer is usually "it depe…
I spent two weeks testing Cursor, GitHub Copilot, Windsurf, and Claude Code on the same set of tasks. Not vibes. Not feature lists. Actual work: bui…
GLM is the new hotness. I'm hearing it from both sides of the AI builder world. Software engineers are talking about it because the benchmark number…
For years, I ran my WordPress sites on OpenLiteSpeed. Fast server, LSCache is genuinely impressive, and the OLS/WordPress combo is hard to beat on r…
How I built a three-task evaluation framework for RNA velocity trajectory inference -- measuring global ordering, pairwise rank preservation, and ro…
CellFateBench is a scientific software and benchmark-engineering project for evaluating reasoning over single-cell genomics workflows. The project w…
Every LLM wire format claims token savings. Nobody proves whether AI models can actually comprehend the format at scale, or produce valid output in…
A blind test across 240 images and 10 professional designers just dropped. Ideogram 4.0 against Gemini 3.1, Grok Imagine, and FLUX.2 Max. The result…
Originally published at deepu.tech . In my release post for LlamaStash I made a claim I need to back up. The wrapper adds zero overhead vs running l…
The Claude Agent SDK exposes three budget tiers ( haiku , sonnet , opus ) and reads its routing target from environment variables on every call. Tha…
Skip to: Full Results | Category Breakdown | The Leaderboard | Methodology TL;DR I built a benchmark suite with 40 vulnerable code patterns across 1…
Book: Prompt Engineering Pocket Guide: Techniques for Getting the Most from LLMs Also by me: Thinking in Go (2-book series) — Complete Guide to Go P…
A research team from the University of Texas at Dallas published LMR-BENCH at EMNLP 2025, asking a specific question: can LLM agents reproduce the c…
A few months ago I shared early results from the A11y LLM Eval project, a benchmark that measures how accessibly LLMs generate UI code. The previous…