Convert speech in audio or video files to text with speaker labels and audio event detection.
Transcription converts spoken content in audio or video files into written text. The AI identifies different speakers, labels them, and tags audio events like music, applause, or laughter.
Attach an audio or video file using the + button in the input bar. Supported formats include MP3, WAV, MP4, MOV, and other common audio and video formats.
Transcribe a meeting recording.
2
Wait for processing
The AI processes the file and generates a text transcript. Longer files take more time.
3
Review the transcript
The agent returns the full text with speaker labels. Each segment is tagged with the speaker who said it.
A transcription result showing speaker-labeled text in the chat.