AI Skills · media · 18 repositories
This page indexes the "media" skill group: 18 open-source repositories. We index only; code is not hosted here and licences belong to the original repositories.
- voicebox — ai', 'cuda', 'mlx', 'qwen3-tts', 'qwen3-tts-ui', 'voice-ai', 'voice-clone', 'whisper
- GPT-SoVITS — text-to-speech', 'tts', 'vits', 'voice-clone', 'voice-cloneai', 'voice-cloning
- VoxCPM — audio', 'deeplearning', 'minicpm', 'multilingual', 'python', 'pytorch', 'speech', 'speech-synthesis
- VoiceStudio — ai', 'audiobook', 'cuda', 'dubbing', 'elevenlabs-alternative', 'huggingface', 'local-first', 'mlx
- VideoLingo — ai-translation', 'dubbing', 'localization', 'video-translation', 'voice-cloning
- Real-Time-Voice-Cloning — https://github.com/CorentinJ/Real-Time-Voice-Cloning
- TTS — deep-learning', 'glow-tts', 'hifigan', 'melgan', 'multi-speaker-tts', 'python', 'pytorch', 'speaker-encoder
- voice-pro — https://github.com/abus-aikorea/voice-pro
- index-tts — bigvgan', 'cross-lingual', 'indextts', 'text-to-speech', 'tts', 'voice-clone', 'zero-shot-tts
- OpenCreator — agent-skills', 'codex-cli', 'deskop-app', 'dubbing', 'image-generation', 'localization', 'skills', 'tts
- OpenVoice — text-to-speech', 'tts', 'voice-clone', 'zero-shot-tts
- MockingBird — ai', 'deep-learning', 'pytorch', 'speech', 'text-to-speech', 'tts
- video-subtitle-remover — https://github.com/YaoFANGUK/video-subtitle-remover
- mlx-audio — apple-silicon', 'audio-processing', 'mlx', 'multimodal', 'speech-recognition', 'speech-synthesis', 'speech-to-text', 'text-to-speech
- SmartSub — deepseek', 'faster-whisper', 'fireredasr', 'funasr', 'ollama', 'openai', 'qwen3-asr', 'sherpa-onnx
- Qwen3-TTS — Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cl
- YouDub-webui — https://github.com/liuzhao1225/YouDub-webui
- MOSS-TTS-Nano — https://github.com/OpenMOSS/MOSS-TTS-Nano