Enable AI Bot Responses to Voice Messages
Overview: AI-Powered Audio Responses
Your CRM's Conversations AI can now process and respond to voice messages and audio files sent by your contacts. When a contact sends audio, the AI automatically transcribes the speech into text, analyzes it using your existing bot training and settings, and delivers an intelligent reply. This feature allows for more natural, conversational interactions where customers can speak instead of type.
Prerequisites
Before you can use this feature, ensure the following are in place:
- Conversations AI is enabled for your sub-account.
- At least one AI Agent is configured and active.
- The messaging channels you intend to use (like WhatsApp or Facebook Messenger) are properly connected.
- The AI Agent is assigned to the specific communication channels where you want it to handle audio.
How to Enable Audio Responses
To configure an AI bot to respond to voice messages, follow these steps:
- Navigate to your sub-account dashboard.
- Go to AI Agents > Conversation AI > Agent List.
- Find the bot you want to configure and click the three-dot menu (⋮) next to its name.
- Select Edit from the menu to open the bot's settings.
- Locate the setting labeled "Also allow this bot to respond to: Voice Notes."
- Toggle this setting to the On position.
- Save your changes.
Once enabled, you can test the feature by sending a voice message from a connected channel like WhatsApp to see if the bot responds appropriately.
Supported Audio Types and Channels
Compatible Audio Formats
The AI bot can process the following types of audio inputs:
- Platform-native voice notes: Recorded directly within apps like WhatsApp, Facebook Messenger, and Instagram.
- Audio file attachments: Supported formats include OGG, MP3, MP4 (audio-only tracks), AAC, M4A, and MPEG. Note that video files (like standard MP4 videos) are not supported for audio transcription.
The bot can also handle multiple audio files sent in quick succession, treating them as part of a single interaction.
Supported Messaging Channels
Audio response functionality is available on channels where Conversations AI is already operational. This includes:
- Facebook Messenger
- Instagram Direct Messages
- SMS/MMS
How Audio Responses Work
Understanding the bot's behavior will help you manage expectations.
- Transcription & Processing: The audio is silently converted to text behind the scenes. This text is then processed by your AI bot using its configured training, prompts, and response mode.
- Response Format: The bot will reply with a standard text message on the channel. It does not generate audio replies.
- Timing and Limits: The bot adheres to your configured Wait Time Before Responding, aggregating multiple messages (including a mix of audio and text) received within that window to craft a single, context-aware reply. It also respects your Maximum Message Limit setting.
- Reviewing Interactions: You can see what the bot "heard" and why it responded the way it did. In the Conversations inbox, open the AI Response Info sidebar for a detailed view of the transcription, the prompt used, and the training sources referenced.
Common Questions
Is there an extra cost for using audio responses?
Usage is billed under your standard Conversations AI usage. Additional messaging fees from the channel provider (like SMS/MMS or WhatsApp) may apply separately.
Can I limit audio handling to specific channels only?
Yes. In the bot's settings, you assign it only to the channels where you want it to be active. The bot will listen and respond to audio only on those assigned channels.
How are multiple voice messages handled?
If a contact sends several audio files within a short period (during your bot's wait time window), they are transcribed and processed together so the bot can generate a single, coherent response based on the full context.
Are there any platform-specific limitations?
Yes. Messaging channels like Facebook Messenger and Instagram have policy windows (e.g., a 24-hour window for messaging) that govern when you can send replies. Your bot's audio response flows must be designed within these constraints.