Voice & Speech

Meilleurs outils IA pour la voix et la parole (2026)

Synthèse vocale, reconnaissance vocale, clonage de voix et IA audio temps réel. De la transcription Whisper à une synthèse vocale digne d'ElevenLabs.

30 outils
Video AI Toolkit — Complete Collection logo

Video AI Toolkit — Complete Collection

Curated video AI tools: Remotion (programmatic video), Manim (math animation), MoviePy (editing), Whisper (speech-to-text), ElevenLabs (voiceover). Build automated video pipelines.

Skill Factory 1,412Skills
F5-TTS — Flow Matching Text-to-Speech logo

F5-TTS — Flow Matching Text-to-Speech

F5-TTS is a diffusion transformer TTS system with flow matching. 14.3K+ GitHub stars. Multi-speaker, voice chat, Gradio UI, CLI inference, 0.04 RTF on L20 GPU. MIT code.

Script Depot 930Skills
Coqui TTS — Deep Learning Text-to-Speech Engine logo

Coqui TTS — Deep Learning Text-to-Speech Engine

Generate speech in 1100+ languages with voice cloning. XTTS v2 streams with under 200ms latency. 44K+ GitHub stars.

TokRepo精选 862Scripts
Fish Speech — Multilingual TTS for 80+ Languages logo

Fish Speech — Multilingual TTS for 80+ Languages

Fish Speech is a state-of-the-art open-source TTS system supporting 80+ languages. 29K+ GitHub stars. 4B dual-AR model, voice cloning, emotional control with 15K+ tags, real-time inference.

AI Open Source 794Skills
GPT-SoVITS — Few-Shot Voice Cloning and Text-to-Speech logo

GPT-SoVITS — Few-Shot Voice Cloning and Text-to-Speech

An open-source TTS system that can clone any voice from just one minute of audio data, combining GPT-style language modeling with VITS synthesis for natural speech generation.

AI Open Source 707Skills
Zonos — Multilingual TTS with Voice Cloning logo

Zonos — Multilingual TTS with Voice Cloning

Zonos is an open-weight TTS model trained on 200K+ hours of speech. 7.2K+ stars. Voice cloning, 5 languages, emotion control. Apache 2.0.

Script Depot 638Scripts
Remotion AI Voiceover Skill — ElevenLabs TTS logo

Remotion AI Voiceover Skill — ElevenLabs TTS

AI skill for adding ElevenLabs text-to-speech voiceover to Remotion videos. Auto-sizes composition duration to match generated audio.

ElevenLabs 619Skills
Index TTS — Industrial Zero-Shot Text-to-Speech System logo

Index TTS — Industrial Zero-Shot Text-to-Speech System

A controllable and efficient zero-shot text-to-speech system built for industrial use, supporting voice cloning and cross-lingual synthesis with high-quality output.

Script Depot 619Skills
Groq Whisper — Sub-Second Speech-to-Text for Voice Agents logo

Groq Whisper — Sub-Second Speech-to-Text for Voice Agents

Whisper-large-v3 on Groq runs 166× realtime — 60-sec clip in <400ms. OpenAI-compat audio.transcriptions endpoint for voice agents.

Groq 614Skills
CosyVoice — Multilingual Voice Generation with LLM-Based TTS logo

CosyVoice — Multilingual Voice Generation with LLM-Based TTS

CosyVoice is an open-source text-to-speech system built on large language models by Alibaba's FunAudioLLM team. It supports 9 languages and 18+ Chinese dialects with zero-shot voice cloning, streaming synthesis, and fine-grained prosody control.

AI Open Source 605Skills
SenseVoice — Multilingual Speech Understanding Model logo

SenseVoice — Multilingual Speech Understanding Model

SenseVoice is an open-source speech foundation model by Alibaba's FunAudioLLM team that performs automatic speech recognition, language identification, speech emotion recognition, and audio event detection in a single model. It supports 50+ languages and runs significantly faster than Whisper.

AI Open Source 560Skills
Deepgram Aura TTS — Text-to-Speech for Voice Agents logo

Deepgram Aura TTS — Text-to-Speech for Voice Agents

Deepgram Aura TTS produces natural English TTS with 250ms TTFA. Streaming WebSocket, 12 voices, tuned for conversational agents not narration.

Deepgram 532Scripts
Dia — Realistic Dialogue Text-to-Speech Model logo

Dia — Realistic Dialogue Text-to-Speech Model

Dia is a 1.6B parameter TTS model by Nari Labs that generates realistic dialogue audio from transcripts. 19.2K+ GitHub stars. Supports multi-speaker dialogue, non-verbal sounds, and voice cloning. Apa

Script Depot 531Skills
Tortoise TTS — Multi-Voice Text-to-Speech Focused on Quality logo

Tortoise TTS — Multi-Voice Text-to-Speech Focused on Quality

A multi-voice TTS system trained with an emphasis on audio quality. Uses autoregressive and diffusion models to produce natural, expressive speech from text.

AI Open Source 464Skills
OmniVoice Studio — Open-Source Voice Cloning and TTS Desktop App logo

OmniVoice Studio — Open-Source Voice Cloning and TTS Desktop App

OmniVoice Studio is a self-hosted desktop application for voice cloning, text-to-speech, dubbing, and dictation. It runs entirely on your local machine, providing a privacy-first alternative to cloud-based voice synthesis services.

Script Depot 355Scripts
Piper — Fast Local Text-to-Speech Engine for 30+ Languages logo

Piper — Fast Local Text-to-Speech Engine for 30+ Languages

Lightweight neural TTS system optimized for Raspberry Pi and edge devices with offline support and dozens of voice models.

AI Open Source 352Configs
VoxCPM — Tokenizer-Free Multilingual Text-to-Speech with Voice Cloning logo

VoxCPM — Tokenizer-Free Multilingual Text-to-Speech with Voice Cloning

Open-source TTS model by OpenBMB that generates natural multilingual speech and clones voices without traditional tokenization.

Script Depot 218Scripts
Index-TTS — Industrial-Grade Zero-Shot Text-to-Speech logo

Index-TTS — Industrial-Grade Zero-Shot Text-to-Speech

An open-source controllable TTS system that clones any voice from a short audio sample, producing natural and expressive speech without fine-tuning.

Script Depot 210Scripts
Silero Models — Pre-Trained Speech AI for STT, TTS, and VAD logo

Silero Models — Pre-Trained Speech AI for STT, TTS, and VAD

Silero Models is a collection of pre-trained enterprise-grade speech models for speech-to-text, text-to-speech, and voice activity detection. The models run efficiently on CPU with PyTorch, requiring no GPU for inference, and support multiple languages out of the box.

AI Open Source 178Configs
EmotiVoice — Multi-Voice Text-to-Speech with Emotion Control logo

EmotiVoice — Multi-Voice Text-to-Speech with Emotion Control

An open-source TTS engine by NetEase Youdao that generates speech with fine-grained emotion control, supporting multiple voices and both Chinese and English with adjustable emotional expression.

AI Open Source 135Configs
whisper.cpp — Local Speech-to-Text in Pure C/C++ logo

whisper.cpp — Local Speech-to-Text in Pure C/C++

High-performance port of OpenAI Whisper in C/C++. No Python, no GPU required. Runs on CPU, Apple Silicon, CUDA, and even Raspberry Pi. Real-time transcription.

Script Depot 3,160代码Skills
Moshi — Real-Time AI Voice Conversation Engine logo

Moshi — Real-Time AI Voice Conversation Engine

Open-source real-time voice AI by Kyutai. Full-duplex speech conversation with 200ms latency, emotion recognition, and on-device processing. Apache 2.0 licensed.

AI Open Source 828Skills
Whisper — OpenAI Speech-to-Text logo

Whisper — OpenAI Speech-to-Text

OpenAI's open-source speech recognition model. Transcribe audio/video to text with word-level timestamps in 99 languages. Essential for subtitle generation.

OpenAI 763Skills
Faster Whisper — 4x Faster Speech-to-Text logo

Faster Whisper — 4x Faster Speech-to-Text

Faster Whisper is a reimplementation of OpenAI Whisper using CTranslate2, up to 4x faster with less memory. 21.8K+ GitHub stars. GPU/CPU, 8-bit quantization, word timestamps, VAD. MIT licensed.

Script Depot 717Skills
LiveKit Agents — Build Real-Time Voice AI Agents logo

LiveKit Agents — Build Real-Time Voice AI Agents

Framework for building real-time voice AI agents. STT, LLM, TTS pipeline with sub-second latency. Supports OpenAI, Anthropic, Deepgram, ElevenLabs. 9.9K+ stars.

LiveKit 688Skills
WhisperX — 70x Faster Speech Recognition logo

WhisperX — 70x Faster Speech Recognition

WhisperX provides 70x realtime speech recognition with word-level timestamps and speaker diarization. 21K+ GitHub stars. Batched inference, under 8GB VRAM. BSD-2-Clause.

Script Depot 681Skills
ChatTTS — Expressive Text-to-Speech for Dialogue logo

ChatTTS — Expressive Text-to-Speech for Dialogue

Generate natural conversational speech with laughter, pauses, and emotion. Optimized for dialogue scenarios. 39K+ GitHub stars.

Script Depot 667Scripts
Remotion Rule: Voiceover logo

Remotion Rule: Voiceover

Remotion skill rule: Adding AI-generated voiceover to Remotion compositions using TTS. Part of the official Remotion Agent Skill for programmatic video in React.

Skill Factory 630Skills
ElevenLabs Python SDK — AI Text-to-Speech logo

ElevenLabs Python SDK — AI Text-to-Speech

Official ElevenLabs Python SDK for AI voice generation. Create realistic voiceovers with 30+ languages, voice cloning, and streaming support.

ElevenLabs 572SkillsCLI Tools
Vapi — Voice AI Agent Platform with STT, LLM & TTS logo

Vapi — Voice AI Agent Platform with STT, LLM & TTS

Vapi glues STT, LLM, TTS, turn-taking into one voice agent API. Build phone agents in minutes. Twilio + Deepgram + ElevenLabs + GPT-4o stack.

Vapi 525Workflows

Les technologies vocales par l'IA

AI Voice Technology

Voice AI has reached a turning point — synthetic speech is now indistinguishable from human narration, and real-time transcription works in 100+ languages. Text-to-Speech (TTS) — ElevenLabs, Coqui TTS, ChatTTS, Fish Speech, and Kokoro generate natural voiceovers with emotional control, multilingual support, and voice cloning from just seconds of sample audio.

Speech-to-Text (STT) — OpenAI's Whisper family (whisper.cpp, WhisperX, Faster Whisper) dominates transcription with near-human accuracy. Self-hosted options run entirely on local hardware for privacy-sensitive applications. Real-Time Voice — Moshi and Dia enable real-time conversational AI with natural turn-taking, interruption handling, and emotional awareness.

Voice Cloning & Synthesis — Clone any voice from a 15-second sample. F5-TTS and Zonos offer open-source voice cloning with quality rivaling commercial APIs. Essential for content creators, podcast producers, and accessibility applications.

Voice is the most natural interface — AI has finally made it programmable.

Questions fréquentes

Quel est le meilleur outil IA de synthèse vocale ?+

Pour la qualité : ElevenLabs mène avec les voix les plus naturelles et le meilleur contrôle émotionnel. Pour l'auto-hébergement : Coqui TTS et Fish Speech offrent une qualité comparable sans coûts d'API. Pour la vitesse : ChatTTS et Kokoro génèrent la parole en temps réel. Pour le multilingue : les pipelines basés sur Whisper combinés à un TTS multilingue gèrent 100+ langues.

L'IA peut-elle cloner ma voix ?+

Oui. Les outils modernes de clonage vocal nécessitent à peine 15 secondes d'échantillon audio. ElevenLabs propose un clonage cloud, tandis que F5-TTS et Zonos offrent des alternatives open source à exécuter localement. La qualité est remarquablement élevée — les voix clonées préservent l'accent, le ton et le style. Obtenez toujours le consentement avant de cloner la voix de quelqu'un.

Quelle est la meilleure reconnaissance vocale open source ?+

Whisper d'OpenAI (via whisper.cpp pour l'inférence locale) est la référence. WhisperX ajoute la diarisation (qui a dit quoi) et les timestamps au mot près. Faster Whisper utilise CTranslate2 pour une vitesse 4x supérieure. Tous fonctionnent localement sans envoyer d'audio à des serveurs externes — critique pour les applications sensibles comme la transcription médicale ou juridique.

Explorer les catégories associées