PPTX → озвученный MP4: как устроен пайплайн, который не ждёт сам себя
Есть презентация и текст лекции — нужен MP4, где каждый слайд висит ровно столько, сколько звучит озвучка. Прототип собирается за час, а дальше начи…
Tech news from the best sources
Есть презентация и текст лекции — нужен MP4, где каждый слайд висит ровно столько, сколько звучит озвучка. Прототип собирается за час, а дальше начи…
The EU AI Act voice watermarking rules took effect on August 2, 2026. Every AI system that generates synthetic audio, image, video or text must now…
В прошлый раз у нас получилось покрыть 20 самых популярных языков России и СНГ с рядом оговорок. Но основным недостатком моделей было то, что значит…
Part of qwen3-tts — a pure C inference engine for Qwen3-TTS. TL;DR Qwen3-TTS ships 9 neutral preset speakers. No emotion control, and cloning a voic…
I benchmarked local voice-cloning models across English, German, Modern Standard Arabic, Spanish, and Mandarin Chinese. Models: OmniVoice int8 Chatt…
Конвейер на Python + Hydra, который превращает папку с аудио в богато размеченный датасет: качество речи, просодия, разборчивость, спикер, транскрип…
A lot of AI apps are starting to mix voice, language models, and generated audio. I built a small Python example that shows that full loop: take an…
Так получилось, что нам посчастливилось принять участие в разработке синтеза для новой версии игры "Ил-2 Штурмовик". Это был длинный путь, но в итог…
Cloud TTS Chirp3-HD with Caching: Fixing Voice Readout for Accessibility As a solo developer, keeping the product lean and accessible is paramount.…
13 лет я тестировала софт, где у бага был адрес: шаг 1, шаг 2, ожидаемый результат, фактический. Нажал — получил. Нажал ещё раз&…
Когда приходишь в Text-to-Speech из классического ML (или даже из CV/NLP), сначала кажется, что всё знакомо: датасет, модель, loss, валидация, поеха…
Мы не так давно опубликовали SAPI5-обёртку для нашего синтеза на 20 языков России и СНГ. В этот раз опять немного сошлись звёзды и мы уже публикуем…
Text-to-speech has gotten good enough that it is no longer just an accessibility feature or a novelty. If you are building an AI app, voice agent, a…
A chapter is the smallest unit a listener actually navigates. They open the audiobook in the middle of Chapter 7, leave it open on the dishes, come…
Typing commands into a serial monitor feels old once you start playing with voice interfaces. So I decided to try something more interesting — build…
My text-to-speech journey started roughly a year ago, when I tried it again and was impressed by how much faster it was than typing. I'd been fascin…
ถ้าคุณใช้ Garudust Agent อยู่แล้ว การเพิ่มความสามารถให้ AI พูดภาษาไทยออกมาเป็นเสียง ทำได้ในไม่กี่ขั้นตอน — ไม่ต้องแก้โค้ดใด ๆ เพราะ Garudust มีระบบ…
Every major voice model lab hosts in us-east or us-west. If you're building voice agents in Sydney, Mumbai, or Jakarta, your audio round-trips to Vi…
How human feedback actually steers TTS fine-tuning Notes on the iteration loop we ran while fine-tuning F5-TTS and StyleTTS2 on a small Northern Eng…
Running modern Python TTS toolchains on non-AVX2 CPUs Notes from getting F5-TTS, StyleTTS2, kokoro/Misaki, and whisper.cpp to work on an AMD Phenom…