Skip to main content
Elumenta supports three types of audio: Text-to-Speech (TTS), Speech-to-Text (STT), and Music Generation, all via the same /api/v2/generate endpoint.

Text-to-Speech

Convert text to natural-sounding speech:

TTS Model Comparison

For real-time applications use elevenlabs-flash. For pre-rendered content (podcasts, audiobooks) use elevenlabs-v2 or openai-tts-hd.

Speech-to-Text

Transcribe audio files:

Music Generation

Two music models for different needs:

MusicGen (Replicate)

ElevenLabs Music

Music Prompt Tips