Tech News
Все новости AI & ML Architecture DevOps Open Source Programming Team Management Testing & QA Web

Последние новости

⚑ Сообщить о проблеме

Tech news from the best sources

Все темы AI Gear News Tech agents ai api architecture automation beginners career database devchallenge devops javascript llm machinelearning mcp opensource performance productivity programming python react security showdev testing tutorial typescript webdev
Все EN RU
EN

A LongMemEval-S number you can reproduce

We held off on posting a benchmark for a long time. Not because we didn't have runs - because most memory benchmarks you read are a number with no w…

aiagentsbenchmarkpython
Dev.to Aug 27, 2026, 21:07 UTC
EN

I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks

I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO m…

aillmmachinelearningbenchmark
Dev.to Aug 27, 2026, 14:09 UTC
EN

We tested "tokenize before you compress" against 452 configurations, and it mostly held up

A few weeks ago my friend @u84u and I ( @ronak-create ) had a simple, slightly annoying question: if LLMs get 30-45% smaller representations of text…

compressionpythonopensourcebenchmark
Dev.to Aug 21, 2026, 01:37 UTC
EN

What should an MCP tool return? I ran 72 trials instead of arguing

There's an argument running about MCP right now. You've probably seen it: a 400-point thread called "MCP is dead?" with real token numbers in it, fo…

mcpllmbenchmarkai
Dev.to Aug 7, 2026, 17:41 UTC
EN

Benchmarking AI Coding Agents on Real Pull Requests

No synthetic puzzles, no contaminated suites: 25 tasks harvested from PRs merged in 2026 across five languages, graded by each project's own held-ou…

aicodingbenchmarkprogramming
Dev.to Aug 1, 2026, 17:25 UTC
EN

AI Daily Digest — August 1, 2026: ARC-AGI-3 Harness Discovery, EU AI Gigafactories, Devin SWE-1.7

🤖💻 AI Daily Digest — August 1, 2026 OpenAI Shows How Two Harness Settings Tripled ARC-AGI-3 Scores OpenAI published a rare technical deep-dive on Ju…

aiagentsbenchmarkhardware
Dev.to Jul 31, 2026, 22:02 UTC
EN

39,000 Torrents: The Bug My Green Benchmark Never Caught

[Torrent911] Inception.2010.TRUEFRENCH.1080p.BluRay.x264-YIFY That's a real filename. The goal: extract "Inception" and "2010" from it, query the TM…

matchingalgorithmsbenchmarkselfhosting
Dev.to Jul 25, 2026, 09:00 UTC
EN

MCPMark v2: InsForge on Sonnet 4.6

Originally published on the InsForge blog , written by Tony Chang (CTO & Co-Founder). Reposted here with permission. In December we published th…

aimcpbenchmarkdatabase
Dev.to Jul 22, 2026, 20:47 UTC
EN

OpenRouter vs Vercel vs LLMGateway Performance

Every AI gateway adds a hop between your app and the model. The question that matters is what that hop costs at the moment your user is staring at a…

aiperformancebenchmarkllm
Dev.to Jul 22, 2026, 17:43 UTC
EN

Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

Frontier models are crushing benchmark after benchmark...so I built a quiz to ask this simple yet humbling question: can a human still beat them? Hu…

aillmhlebenchmark
Dev.to Jul 20, 2026, 23:05 UTC
EN

The Same RTX 5090, but the GPU Sat Idle — a CPU-Bound Go Solver and the Case for L2 Cache

This is the second in a short series that benchmarks a single RTX 5090 by re-running published Go solvers — programs that don't just play Go but pro…

cpugpubenchmarkhardware
Dev.to Jul 18, 2026, 05:08 UTC
EN

One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

You don't need to know anything about Go to read this. The game is just the fixed yardstick. The story is a hardware benchmark: the same program, th…

gpubenchmarkmachinelearningcuda
Dev.to Jul 18, 2026, 00:33 UTC
EN

Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano

Round 7 ended on a cliffhanger I couldn't stop thinking about. Qwen 3.6 35B-A3B built the entire feature — read the codebase, wrote the files, got a…

modelshowdownbenchmarkaillm
Dev.to Jul 15, 2026, 21:52 UTC
EN

AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

What Changed Large language models (LLMs) have demonstrated proficiency in high-school and olympiad-style mathematics. However, their performance in…

llmmathematicsbenchmarkproofgeneration
Dev.to Jul 14, 2026, 11:20 UTC
EN

Which LLM should I actually code with? I built a small benchmark to find out

LLM code benchmark — Peculiar Engineer A small, self-run coding benchmark: 3 models on 14 problems across 3 languages, scored on pass@k, cost, and s…

aillmbenchmarkprogramming
Dev.to Jul 12, 2026, 03:41 UTC
EN

I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

I run ezyZip , a browser-based archive tool, so "which format should I use?" is a question I field constantly. The honest answer is usually "it depe…

compressionzipbenchmarkcli
Dev.to Jul 10, 2026, 15:41 UTC
EN

AI Coding Tools Benchmark 2026: Cursor vs Copilot vs Windsurf vs Claude Code

I spent two weeks testing Cursor, GitHub Copilot, Windsurf, and Claude Code on the same set of tasks. Not vibes. Not feature lists. Actual work: bui…

codingbenchmarkcursorgithubcopilot
Dev.to Jul 7, 2026, 01:00 UTC
EN

I built a neutral benchmarking layer for quantum simulators in Rust — and it revealed a silent disagreement between two backends

rustquantumcomputingopensourcebenchmark
Dev.to Jul 4, 2026, 18:26 UTC
EN

GLM Is the New Hotness, So Let's Test It On the Homelab

GLM is the new hotness. I'm hearing it from both sides of the AI builder world. Software engineers are talking about it because the benchmark number…

modelshowdownbenchmarkaillm
Dev.to Jun 30, 2026, 20:14 UTC
EN

Too cheap to be good? Think again.

For years, I ran my WordPress sites on OpenLiteSpeed. Fast server, LSCache is genuinely impressive, and the OLS/WordPress combo is hard to beat on r…

aibenchmarkdevopswebdev
Dev.to Jun 23, 2026, 20:38 UTC
EN

A UMAP With Arrows Is Not a Benchmark. This Is

How I built a three-task evaluation framework for RNA velocity trajectory inference -- measuring global ordering, pairwise rank preservation, and ro…

benchmarkbioinformaticsrnascientificsoftware
Dev.to Jun 16, 2026, 23:51 UTC
EN

Engineering CellFateBench: A Reproducible Python Benchmark for Single-Cell Genomics Reasoning

CellFateBench is a scientific software and benchmark-engineering project for evaluating reasoning over single-cell genomics workflows. The project w…

bioinformaticsgenomicsbenchmarkpython
Dev.to Jun 16, 2026, 22:14 UTC
EN

LLM Wire Format Benchmark: Which Format Can AI Actually Read and Write?

Every LLM wire format claims token savings. Nobody proves whether AI models can actually comprehend the format at scale, or produce valid output in…

llmbenchmarkaiwebdev
Dev.to Jun 7, 2026, 00:11 UTC
EN

Ideogram 4.0 is Good. Just Good.

A blind test across 240 images and 10 professional designers just dropped. Ideogram 4.0 against Gemini 3.1, Grok Imagine, and FLUX.2 Max. The result…

aireviewimagegenerationbenchmark
Dev.to Jun 6, 2026, 06:24 UTC
EN

How fast is LlamaStash? Overhead, throughput, and a fair comparison with Ollama and LM Studio

Originally published at deepu.tech . In my release post for LlamaStash I made a claim I need to back up. The wrapper adds zero overhead vs running l…

aillamacppbenchmarkllm
Dev.to Jun 2, 2026, 11:34 UTC
EN

Benchmarking the Claude Agent SDK on a local LLM: Haiku and Sonnet tier performance

The Claude Agent SDK exposes three budget tiers ( haiku , sonnet , opus ) and reads its routing target from environment variables on every call. Tha…

llmclaudellamacppbenchmark
Dev.to May 28, 2026, 08:31 UTC
EN

I Benchmarked 17 ESLint Security Plugins. Only One Found Every Vulnerability.

Skip to: Full Results | Category Breakdown | The Leaderboard | Methodology TL;DR I built a benchmark suite with 40 vulnerable code patterns across 1…

securityeslintjavascriptbenchmark
Dev.to May 25, 2026, 14:27 UTC
EN

Multi-Shot vs Zero-Shot: When Adding Examples Actually Hurts Accuracy

Book: Prompt Engineering Pocket Guide: Techniques for Getting the Most from LLMs Also by me: Thinking in Go (2-book series) — Complete Guide to Go P…

aillmpromptbenchmark
Dev.to May 24, 2026, 09:34 UTC
EN

LMR-BENCH: Can LLM Agents Reproduce NLP Research Code? (EMNLP 2025)

A research team from the University of Texas at Dallas published LMR-BENCH at EMNLP 2025, asking a specific question: can LLM agents reproduce the c…

benchmarkresearchreproducibilityllmagentspaperpoc
Dev.to May 22, 2026, 12:16 UTC
EN

AI-generated accessibility, an update — frontier models still fail, but skills change the game

A few months ago I shared early results from the A11y LLM Eval project, a benchmark that measures how accessibly LLMs generate UI code. The previous…

a11yllmaibenchmark
Dev.to May 21, 2026, 14:40 UTC

© Tech News — Агрегатор новостей

English Русский
Карта сайта Правовая информация Конфиденциальность Условия использования Авторские права / Удаление Контакт DSA

Выход с сайта

Вы собираетесь открыть внешний сайт:

Продолжить →