ElevenLabs
拟真度极高的语音合成
拟真度极高的语音合成
科大讯飞的语音合成服务
AI 配音配乐工具,多情绪音色与字幕时间轴
A scalable generative AI framework built for resear…
State-of-the-Art Deep Learning scripts organized by…
Towards Human-Sounding Speech
Generate audiobooks from EPUBs, PDFs and text with …
中文声音克隆与 TTS
文本生成完整歌曲
AI 音乐生成与混音
字节出品的 AI 音乐生成
The Generative AI Landscape - A Collection of Aweso…
Lab Materials for MIT 6.S191: Introduction to Deep …
🤗 Transformers: the model-definition framework for…
开源多语种语音识别
语音转写与会议纪要
Privacy first, AI meeting assistant with 4x faster …
Drench yourself in Deep Learning, Reinforcement Lea…
Swap GPT for any LLM by changing a single line of c…
Quantization, kernels, runtime and inference engine…
The python library for real-time communication
The media player for language learning, with dual s…
MOSS‑TTS Family is an open‑source speech and sound …
Hold a key, speak, release — AI-polished text appea…
Implementation of MusicLM, Google's new SOTA model …
A single Gradio + React WebUI with extensions for A…
The official Python SDK for the ElevenLabs API.
🔊 Awesome list for Whisper — an open-source AI-pow…
An open-source ChatGPT app with a voice
AI-powered cross-platform e-book reader with semant…
👻 Proxy API gateway for Kiro IDE & CLI (Amazon Q D…
Instant, controllable, local pre-trained AI models …
ASR/STT subtitle generator. Uses Qwen3-ASR, local L…
Let Claude (or any LLM) actually watch a video — sc…
Audio generation using diffusion models, in PyTorch.
Free, high-quality text-to-speech API endpoint to r…
Open-source industrial-grade ASR models supporting …
Text-To-Speech, RAG, and LLMs. All local!
A timeline of the latest AI models for audio genera…
OpenAI API client for Kotlin with multiplatform and…
百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R…
A high-performance inference engine for AI models
Implementation of SoundStorm, Efficient Parallel Au…
Talk to your Mac, query your docs, no cloud require…
List of Machine Learning, AI, NLP solutions for iOS…
the open-source virtual assistant for Ubuntu based …
Implementation of Natural Speech 2, Zero-shot Speec…
A webui for different audio related Neural Networks
SincNet is a neural architecture for efficiently pr…
A practical lab for building, testing, and evaluati…
Finding the genre of a song with Deep Learning
AI Audio Datasets (AI-ADS) 🎵, including Speech, Mu…
Turn an epub or text file into an audiobook
Implementation of Band Split Roformer, SOTA Attenti…
a list of demo websites for automatic music generat…
:robot::art::guitar:A curated list of awesome proje…
High-performance Text-to-Speech server with OpenAI-…
Implementation of Voicebox, new SOTA Text-to-speech…
Fast voice assistant powered by Groq, Cartesia, and…
Implementation of MusicLM, a text to music model pu…
Implementation of E2-TTS, "Embarrassingly Easy Full…
A curated compilation of AI-driven generative music…
Open-source AI voice typing for macOS, Windows, and…
Audio Development Tools (ADT) is a project for adva…
The official JavaScript (Node) library for the Elev…
A feature-rich portal to chat with GPT-4, Claude, G…
Free tool to create viral videos from YouTube, gene…
Music discovery tool that integrates with Lidarr an…
[LREC-COLING 2024 (Oral), Interspeech 2024 (Oral), …
VoxNovel: generate audiobooks giving each character…
Voice Activity Detection based on Deep Learning & T…
UTokyo-SaruLab MOS Prediction System
AIUI is a platform enabling seamless two-way verbal…
infinifi plays gentle lofi music in the background …
Speech-to-speech AI assistant with natural conversa…
A ComfyUI custom node suite for Qwen3-TTS, supporti…
Tracking the progress in non-autoregressive generat…
Snap any video URL or audio file into plaintext. No…
Tegridy MIDI Dataset for precise and effective Musi…
PlayHT Python SDK - AI Text-to-Speech Streaming & V…