FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

speech-ai · GitHub Topics · GitHub

#

speech-ai

Here are 57 public repositories matching this topic...

Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.

  • Updated Aug 28, 2026
  • Go

A New End-to-end Framework for Evaluating Voice Agents

  • Updated Aug 27, 2026
  • Python

Project for speech bubble

  • Updated Aug 15, 2025
  • Python

🇺🇦 Open Source Ukrainian Text-to-Speech datasets

  • Updated Feb 24, 2025
  • Python

A Docker-based OpenAI-compatible Text-to-Speech API server powered by Kyutai's TTS models with GPU acceleration support.

  • Updated Jul 12, 2025
  • Python

Just a simple multimodal avatar interaction platform

  • Updated Jul 30, 2026
  • JavaScript

A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.

  • Updated Jul 14, 2025
  • TypeScript

A unified benchmarking framework for evaluating Voice AI agents across conversational quality, audio realism, latency metrics, and safety guardrails with scalable multi-language stress testing.

  • Updated Feb 26, 2026
  • Python

MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.

  • Updated Aug 28, 2026
  • Python

Open-source real-time Voice AI infrastructure in Go. Stream audio via WebRTC or WebSocket, connect STT → LLM → TTS pipelines, and build scalable voice agents and conversational AI applications.

  • Updated Jun 15, 2026
  • Go

🇺🇦 Ukrainian RAD-TTS++ models (decoder + models with 3 voices) and HiFiGAN model

  • Updated Feb 27, 2025

A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.

  • Updated Jun 24, 2025
  • TypeScript

A source-linked directory of free and trial LLM APIs, multimodal models, embeddings, speech, translation, safety, and other inference endpoints. Companion catalog for freellmapi.io.

  • Updated Aug 14, 2026

Legacy Speech AI examples with migration links to the current Brainiall TTS and transcription services.

  • Updated Aug 3, 2026
  • JavaScript

A curated list of the best Text-to-Speech, speech synthesis, and voice-cloning research — models, papers, benchmarks, and toolkits, focused on 2025–2026.

  • Updated Jun 28, 2026

中文 ASR 评测工具箱 · micro-CER 对比 FunASR/Whisper/llama.cpp · 一条命令出报告 · 自带迷你测试集 · Mandarin ASR benchmark toolkit

  • Updated Jul 6, 2026
  • Python

Multilingual AI speech studio for Text-to-Speech, Speech-to-Text, Voice Cloning, and Audio Enhancement in Uzbek, English, and Korean.

  • Updated Aug 28, 2026
  • Python

Code-switching ASR adaptation for strong multilingual speech recognition models. Synthetic CSW data generation, Whisper adaptation, Bayesian LoRA (BLoRA), and robust multilingual ASR evaluation.

  • Updated Jul 9, 2026

Interactive documentation helper for Sarvam AI APIs — grounded answers from official docs for multilingual speech & language products

  • Updated Jul 20, 2026
  • Python

Improve this page

Add a description, image, and links to the speech-ai topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the speech-ai topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL