Как я ускорил TypeScript-типы в 15.7 раза
Иногда TypeScript думает мучительно долго. У меня это вылезло не самым приятным способом: @_chenglou — он работал над React, ReasonML и ReScript, а…
Tech news from the best sources
Иногда TypeScript думает мучительно долго. У меня это вылезло не самым приятным способом: @_chenglou — он работал над React, ReasonML и ReScript, а…
Code: Megapixel99/capture-the-flag In April I ran five games of an AI capture-the-flag tournament between five small open-weight models (1.0B to 2.5…
Our CEO Rob Imbeault published a piece on LinkedIn this week about a result our team posted: 99.95% on LoCoMo , the most cited benchmark for long-te…
Recently, I've added a bunch of hype-monsters to my AI Werewolf : Kimi K3 Qwen 3.8 Max, Qwen 3.7 Plus, Qwen 3.7 Flash MiniMax M3 Plus the ones I've…
.NET Matrix — это новый открытый проект, который сравнивает .NET-библиотеки внутри одной категории по трём аспектам: возможности , скорость и исполь…
Five stories moved the AI-coding world today. None are about a single model winning forever — they are about the ground shifting under who runs the…
Video AI systems consistently fail to track what happens when the camera looks away: when a scene pans away from an object in motion and returns, cu…
A large-scale audit of AI-as-judge evaluation — covering over half a million individual judgments — finds that AI judges are consistently reliable b…
The headline number is 95% on SWE-bench Verified. That's the score attached to Claude Fable 5, Anthropic's new general-access model in the Mythos cl…
I wrote about hardware benchmarks twice this week. Different problem this time. Same machines. I have a Mac for daily work, a Linux box that runs a…
I wrote a post yesterday about why GPUs barely help small text embeddings at batch=1. Different workload, same machines. This time I ran a local LLM…
Most "just use a GPU" advice is wrong for how anyone actually runs small models. I spent yesterday benchmarking a 33M parameter embedding model acro…
Overview PDF text extraction is a common pre-processing step in data pipelines — ingesting research papers, legal documents, or reports before embed…
Overview PageRank is the canonical graph algorithm. NetworkX implements it in pure Python — its dict-of-dict adjacency representation means every po…
Author: Tobie Morgan Hitchcock One engine, multi-workloads, full durability. You can explore the full results, methodology, and per-database breakdo…
We added a npm run ilb:flagship:smoke gate to the quality script. It's small: for each flagship rule with a labeled corpus, run the rule against vul…
By Natnael Alemseged The gap that τ²-Bench retail cannot measure Tenacious is a B2B sales automation company. Its agent produces outreach emails for…