Tech News
Все новости AI & ML Architecture DevOps Open Source Programming Team Management Testing & QA Web

Последние новости

⚑ Сообщить о проблеме

Tech news from the best sources

Все темы - игры AI Gear News Tech agents ai api architecture automation beginners career database devops javascript llm machinelearning mcp opensource performance productivity programming python react security showdev testing tutorial typescript webdev
Все EN RU
RU

Как я ускорил TypeScript-типы в 15.7 раза

Иногда TypeScript думает мучительно долго. У меня это вылезло не самым приятным способом: @_chenglou — он работал над React, ReasonML и ReScript, а…

typescriptperformancetestsbenchmarks
Habr Aug 17, 2026, 07:03 UTC
EN

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Code: Megapixel99/capture-the-flag In April I ran five games of an AI capture-the-flag tournament between five small open-weight models (1.0B to 2.5…

llmmeasurementsecuritybenchmarks
Dev.to Aug 15, 2026, 06:51 UTC
EN

We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

Our CEO Rob Imbeault published a piece on LinkedIn this week about a result our team posted: 99.95% on LoCoMo , the most cited benchmark for long-te…

aibenchmarkswebdevsoftware
Dev.to Aug 12, 2026, 02:12 UTC
EN

How I tried to write an article about slow Chinese LLMs

Recently, I've added a bunch of hype-monsters to my AI Werewolf : Kimi K3 Qwen 3.8 Max, Qwen 3.7 Plus, Qwen 3.7 Flash MiniMax M3 Plus the ones I've…

aillmbenchmarkslatency
Dev.to Aug 6, 2026, 15:48 UTC
RU

.NET Matrix: взвешенный выбор библиотеки, а не по звёздам на GitHub

.NET Matrix — это новый открытый проект, который сравнивает .NET-библиотеки внутри одной категории по трём аспектам: возможности , скорость и исполь…

.netlibraryframeworkbenchmarksbenchmarkbenchmarkingbenchmarkdotnet
Habr Jul 30, 2026, 14:11 UTC
EN

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Five stories moved the AI-coding world today. None are about a single model winning forever — they are about the ground shifting under who runs the…

newsindustrybenchmarkslaunches
Dev.to Jul 11, 2026, 21:24 UTC
EN

Turn the camera away, and the AI's world freezes

Video AI systems consistently fail to track what happens when the camera looks away: when a scene pans away from an object in motion and returns, cu…

worldmodelsvideogenerationroboticsbenchmarks
Dev.to Jul 2, 2026, 00:17 UTC
EN

Reliable, and still wrong

A large-scale audit of AI-as-judge evaluation — covering over half a million individual judgments — finds that AI judges are consistently reliable b…

evaluationllmasjudgebenchmarks
Dev.to Jul 1, 2026, 21:43 UTC
EN

Claude Fable 5 Scores 95% on SWE-bench, Then Hands Off to Opus 4.8

The headline number is 95% on SWE-bench Verified. That's the score attached to Claude Fable 5, Anthropic's new general-access model in the Mythos cl…

anthropicclaudebenchmarkssafety
Dev.to Jun 12, 2026, 08:21 UTC
EN

Cross-Machine Memory Query: About 20 Milliseconds, Most Days

I wrote about hardware benchmarks twice this week. Different problem this time. Same machines. I have a Mac for daily work, a Linux box that runs a…

performancebenchmarksmachinelearningwireguard
Dev.to Jun 3, 2026, 14:26 UTC
EN

An AMD GPU Beat My Mac on Llama 8B. The Same GPU Lost on Phi-3.

I wrote a post yesterday about why GPUs barely help small text embeddings at batch=1. Different workload, same machines. This time I ran a local LLM…

performancebenchmarksmachinelearninggpu
Dev.to Jun 2, 2026, 18:28 UTC
EN

Your GPU Probably Isn't Helping Your Retrieval System

Most "just use a GPU" advice is wrong for how anyone actually runs small models. I spent yesterday benchmarking a 33M parameter embedding model acro…

performancebenchmarksmachinelearninggpu
Dev.to Jun 2, 2026, 16:28 UTC
EN

pypdf vs PdfPig: Text Extraction at Scale

Overview PDF text extraction is a common pre-processing step in data pipelines — ingesting research papers, legal documents, or reports before embed…

dotnetcsharpperformancebenchmarks
Dev.to May 31, 2026, 18:20 UTC
EN

NetworkX vs CSR + TensorPrimitives: PageRank on 28M Edges

Overview PageRank is the canonical graph algorithm. NetworkX implements it in pure Python — its dict-of-dict adjacency representation means every po…

dotnetcsharpperformancebenchmarks
Dev.to May 31, 2026, 18:20 UTC
EN

SurrealDB 3.x by the numbers

Author: Tobie Morgan Hitchcock One engine, multi-workloads, full durability. You can explore the full results, methodology, and per-database breakdo…

surrealdbdatabasebenchmarksnews
Dev.to May 29, 2026, 19:20 UTC
EN

What ground truth caught that unit tests missed: 3 real bugs in 9 flagship lint rules

We added a npm run ilb:flagship:smoke gate to the quality script. It's small: for each flagship rule with a labeled corpus, run the rule against vul…

staticanalysiseslinttestingbenchmarks
Dev.to May 14, 2026, 05:18 UTC
EN

When Generic Benchmarks Fail: Building a Sales-Domain Evaluation Bench from Scratch

By Natnael Alemseged The gap that τ²-Bench retail cannot measure Tenacious is a B2B sales automation company. Its agent produces outreach emails for…

machinelearningllmbenchmarksai
Dev.to May 2, 2026, 18:16 UTC

© Tech News — Агрегатор новостей

English Русский
Карта сайта Правовая информация Конфиденциальность Условия использования Авторские права / Удаление Контакт DSA

Выход с сайта

Вы собираетесь открыть внешний сайт:

Продолжить →