Skip to main content
ElevenLabs Documentation Docs

Search documentation

Type to search this documentation.

On this pageOverview

Multichannel speech-to-text

The multichannel Speech to Text feature enables you to transcribe audio files where each channel contains a distinct speaker. This is particularly useful for recordings where speakers are isolated on separate audio channels, providing cleaner transcriptions without the need for speaker diarization.

Each channel is processed independently and automatically assigned a speaker ID based on its channel number (channel 0 → speaker_0, channel 1 → speaker_1, etc.). The system extracts individual channels from your input audio file and transcribes them in parallel. By default the API returns one transcript per channel; set multichannel_output_style=combined to instead receive a single transcript with all channels merged into one list sorted by start time, with each word tagged by its channel_index.

  • Stereo interview recordings - Interviewer on left channel, interviewee on right channel
  • Multi-track podcast recordings - Each participant recorded on a separate track
  • Call center recordings - Agent and customer separated on different channels
  • Conference recordings - Individual participants isolated on separate channels
  • Court proceedings - Multiple parties recorded on distinct channels
  • An ElevenLabs account with an API key
  • Multichannel audio file (WAV, MP3, or other supported formats)
  • Maximum 5 channels per audio file
  • Each channel should contain only one speaker

Ensure your audio file has speakers isolated on separate channels. The multichannel feature supports up to 5 channels, with each channel mapped to a specific speaker:

  • Channel 0 → speaker_0
  • Channel 1 → speaker_1
  • Channel 2 → speaker_2
  • Channel 3 → speaker_3
  • Channel 4 → speaker_4

When making a speech-to-text request, you must set:

  • use_multi_channel: true
  • diarize: false (multichannel mode handles speaker separation via channels)

Optionally, control the response shape with:

  • multichannel_output_style: separate (default) returns one transcript per channel. combined merges all channels into a single transcript whose words are sorted by start time, each carrying a channel_index — matching the standard single-channel response shape. combined requires timestamps (timestamps_granularity must not be none) and is not supported with webhook delivery or entity detection/redaction.

The num_speakers parameter cannot be used with multichannel mode as the speaker count is automatically determined by the number of channels. Multichannel mode assumes there will exactly one speaker per channel. If there are more, it will assign the same speaker id to all speakers in the channel.

By default (multichannel_output_style=separate), multichannel audio returns a different response format than single-channel:

Single channel response

Single channel response
{
  "language_code": "en",
  "language_probability": 0.98,
  "text": "Hello world",
  "words": [...]
}

Multichannel response

Multichannel response
{
  "transcripts": [
    {
      "language_code": "en",
      "language_probability": 0.98,
      "text": "Hello from channel one.",
      "channel_index": 0,
      "words": [...]
    },
    {
      "language_code": "en",
      "language_probability": 0.97,
      "text": "Greetings from channel two.",
      "channel_index": 1,
      "words": [...]
    }
  ]
}

Combined response (multichannel_output_style=combined)

Combined response (multichannel_output_style=combined)
{
  "language_code": "en",
  "language_probability": 0.98,
  "text": "Hello from channel one. Greetings from channel two.",
  "words": [
    { "text": "Hello", "start": 0.0, "end": 0.5, "type": "word", "speaker_id": "speaker_0", "channel_index": 0 },
    { "text": "Greetings", "start": 0.6, "end": 1.2, "type": "word", "speaker_id": "speaker_1", "channel_index": 1 }
  ]
}

With multichannel_output_style=combined, the response uses the same flat shape as a single-channel transcription (top-level text and words, no transcripts array), with all channels merged into one list sorted by start time. Every word includes a channel_index (and speaker_id) identifying its channel.

Here's a complete example of transcribing a stereo audio file with two speakers:

Python

Python
from elevenlabs import ElevenLabs

elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")

def transcribe_multichannel(audio_file_path):
    with open(audio_file_path, 'rb') as audio_file:
        result = elevenlabs.speech_to_text.convert(
            file=audio_file,
            model_id='scribe_v2',
            use_multi_channel=True,
            diarize=False,
            timestamps_granularity='word'
        )
    return result

# Process the response

result = transcribe_multichannel('stereo_interview.wav')

if hasattr(result, 'transcripts'): # Multichannel response
    for transcript in result.transcripts:
        channel = transcript.channel_index
        text = transcript.text
        print(f"Channel {channel} (speaker_{channel}): {text}")
    else: # Single channel response (fallback)
        print(f"Text: {result.text}")

JavaScript

JavaScript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import fs from "fs";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

async function transcribeMultichannel(audioFilePath) {
  try {
    const audioFile = fs.createReadStream(audioFilePath);

    const result = await elevenlabs.speechToText.convert({
      file: audioFile,
      modelId: "scribe_v2",
      useMultiChannel: true,
      diarize: false,
      timestampsGranularity: "word",
    });

    // With useMultiChannel: true the SDK returns one transcript per channel
    result.transcripts.forEach((transcript) => {
      console.log(`Channel ${transcript.channelIndex}: ${transcript.text}`);
    });

    return result;
  } catch (error) {
    console.error("Error transcribing audio:", error);
    throw error;
  }
}

cURL

cURL
curl -X POST "https://api.elevenlabs.io/v1/speech-to-text" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "file=@stereo_audio_file.wav" \
  -F "model_id=scribe_v2" \
  -F "use_multi_channel=true" \
  -F "diarize=false" \
  -F "timestamps_granularity=word"

The easiest way to get a time-ordered, conversation-style transcript is to request multichannel_output_style=combined — the API returns a single words list, already sorted by start time, with a channel_index and speaker_id on each word:

Combined output (recommended)

Combined output (recommended)
with open("stereo_interview.wav", "rb") as audio_file:
    result = elevenlabs.speech_to_text.convert(
        file=audio_file,
        model_id="scribe_v2",
        use_multi_channel=True,
        multichannel_output_style="combined",
        diarize=False,
        timestamps_granularity="word",
    )

for word in result.words:
    if word.type == "word":
        print(f"speaker_{word.channel_index}: {word.text}")

If you're using the default separate output, you can merge the per-channel transcripts client-side instead:

Python
def create_conversation_transcript(multichannel_result):
    """Create a conversation-style transcript with speaker labels"""
    all_words = []

    if hasattr(multichannel_result, 'transcripts'):
        # Collect all words from all channels
        for transcript in multichannel_result.transcripts:
            for word in transcript.words or []:
                if word.type == 'word':
                    all_words.append({
                        'text': word.text,
                        'start': word.start,
                        'speaker_id': word.speaker_id,
                        'channel': transcript.channel_index
                    })

    # Sort by timestamp
    all_words.sort(key=lambda w: w['start'])

    # Group consecutive words by speaker
    conversation = []
    current_speaker = None
    current_text = []

    for word in all_words:
        if word['speaker_id'] != current_speaker:
            if current_text:
                conversation.append({
                    'speaker': current_speaker,
                    'text': ' '.join(current_text)
                })
            current_speaker = word['speaker_id']
            current_text = [word['text']]
        else:
            current_text.append(word['text'])

    # Add the last segment
    if current_text:
        conversation.append({
            'speaker': current_speaker,
            'text': ' '.join(current_text)
        })

    return conversation

# Format the output
conversation = create_conversation_transcript(result)
for turn in conversation:
    print(f"{turn['speaker']}: {turn['text']}")

Multichannel transcription supports webhook delivery for asynchronous processing:

Python
from elevenlabs import ElevenLabs

elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")

async def transcribe_multichannel_with_webhook(audio_file_path):
    with open(audio_file_path, 'rb') as audio_file:
        result = await elevenlabs.speech_to_text.convert_async(
            file=audio_file,
            model_id='scribe_v2',
            use_multi_channel=True,
            diarize=False,
            webhook=True  # Enable webhook delivery
        )

    print(f"Transcription started with task ID: {result.task_id}")
    return result.task_id

Error: Multichannel mode does not support diarization and assigns speakers based on the channel they speak on.

Solution: Always set diarize=false when using multichannel mode.

Error: Cannot specify num_speakers when use_multi_channel is enabled. The number of speakers is automatically determined by the number of channels. Solution: Remove the num_speakers parameter from your request.

Error: Multichannel mode supports up to 5 channels, but the audio file contains X channels.

Solution: Process only the first 5 channels or pre-process your audio to reduce channel count.

Error: multichannel_output_style='combined' requires timestamps; set timestamps_granularity to 'word' or 'character'.

Solution: Combined output sorts words by time, so set timestamps_granularity to word (the default) or character.

Error: multichannel_output_style='combined' is not yet supported with webhook delivery.

Solution: Use a synchronous request with combined, or keep the default separate output when using webhooks and merge client-side.

The concurrency cost increases linearly with the number of channels. A 60-second 3-channel file has 3x the concurrency cost of a single-channel file.

You can estimate the processing time for multichannel audio using the following formula:

Processing Time=(D⋅0.3)+2+(N⋅0.5)Processing\ Time = (D \cdot 0.3) + 2 + (N \cdot 0.5)

Where:

  • $D$ = file duration in seconds
  • $N$ = number of channels
  • $0.3$ = processing speed factor (approximately 30% of real-time)
  • $2$ = fixed overhead in seconds
  • $0.5$ = per-channel overhead in seconds

Example: For a 60-second stereo file (2 channels):

Processing Time=(60⋅0.3)+2+(2⋅0.5)=18+2+1=21 secondsProcessing\ Time = (60 \cdot 0.3) + 2 + (2 \cdot 0.5) = 18 + 2 + 1 = 21\ seconds

For large multichannel files, consider streaming or chunking:

Python

Python
def process_large_multichannel_file(file_path, chunk_duration=300):
    """Process large files in chunks (5-minute segments)"""

    from pydub import AudioSegment
    from elevenlabs import ElevenLabs
    import os

    elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")
    audio = AudioSegment.from_file(file_path)
    duration_ms = len(audio)
    chunk_size_ms = chunk_duration * 1000

    all_transcripts = []

    for start_ms in range(0, duration_ms, chunk_size_ms):
        end_ms = min(start_ms + chunk_size_ms, duration_ms)

        # Extract chunk
        chunk = audio[start_ms:end_ms]
        chunk_file = f"temp_chunk_{start_ms}.wav"
        chunk.export(chunk_file, format="wav")

        # Transcribe chunk using SDK
        with open(chunk_file, 'rb') as audio_file:
            result = elevenlabs.speech_to_text.convert(
                file=audio_file,
                model_id='scribe_v2',
                use_multi_channel=True,
                diarize=False,
                timestamps_granularity='word'
            )

        # Adjust timestamps
        if hasattr(result, 'transcripts'):
            for transcript in result.transcripts:
                for word in transcript.words or []:
                    word.start += start_ms / 1000
                    word.end += start_ms / 1000
            all_transcripts.extend(result.transcripts)

        # Clean up
        os.remove(chunk_file)

    return all_transcripts

JavaScript

JavaScript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { exec } from "child_process";
import fs from "fs";
import path from "path";
import { promisify } from "util";

const execAsync = promisify(exec);
const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

async function processLargeMultichannelFile(filePath, chunkDuration = 300) {
  /**
   * Process large files in chunks (5-minute segments)
   * Requires ffmpeg to be installed
   */

  // Get audio duration using ffprobe
  const { stdout } = await execAsync(
    `ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 "${filePath}"`
  );
  const durationSeconds = parseFloat(stdout);

  const allTranscripts = [];

  for (let startSeconds = 0; startSeconds < durationSeconds; startSeconds += chunkDuration) {
    const endSeconds = Math.min(startSeconds + chunkDuration, durationSeconds);
    const chunkFile = path.join(path.dirname(filePath), `temp_chunk_${startSeconds}.wav`);

    // Extract chunk using ffmpeg
    await execAsync(
      `ffmpeg -i "${filePath}" -ss ${startSeconds} -t ${chunkDuration} -c:a pcm_s16le "${chunkFile}" -y`
    );

    try {
      // Transcribe chunk
      const audioFile = fs.createReadStream(chunkFile);
      const result = await elevenlabs.speechToText.convert({
        file: audioFile,
        modelId: "scribe_v2",
        useMultiChannel: true,
        diarize: false,
        timestampsGranularity: "word",
      });

      // Adjust timestamps
      if (result.transcripts) {
        for (const transcript of result.transcripts) {
          for (const word of transcript.words || []) {
            word.start += startSeconds;
            word.end += startSeconds;
          }
        }
        allTranscripts.push(...result.transcripts);
      }
    } finally {
      // Clean up
      fs.unlinkSync(chunkFile);
    }
  }

  return { transcripts: allTranscripts };
}

What happens if my audio has more than 5 channels?

Section titled “What happens if my audio has more than 5 channels?”

The API will return an error. You'll need to either select which 5 channels to send to the API or mix down some channels before sending them to the API.

Can I process mono audio with multichannel mode?

Section titled “Can I process mono audio with multichannel mode?”

Yes, but it's unnecessary. If you send mono audio with use_multi_channel=true, you'll receive a standard single-channel response, not the multichannel format.

Can I get one combined transcript instead of separate per-channel transcripts?

Section titled “Can I get one combined transcript instead of separate per-channel transcripts?”

Yes. Set multichannel_output_style=combined to receive a single transcript with all channels merged and sorted by start time, each word tagged with its channel_index. This matches the standard single-channel response shape. It requires timestamps and isn't available with webhook delivery.

Speaker IDs are deterministic based on channel number: channel 0 becomes speaker_0, channel 1 becomes speaker_1, and so on.

Yes, each channel is processed independently and can detect different languages. The language detection happens per channel. With multichannel_output_style=combined, the top-level language_code reflects the most confident channel, while each word still carries its channel_index.

Full Speech to Text API reference and parameters.

Receive transcription results asynchronously via webhook.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu