Captioning and Transcripts
Captioning and Transcripts
Captioning
Captioning video content is an essential part of creating high-quality educational resources that improve student outcomes and user engagement. Captions and transcripts provide a text alternative to audio, offering viewers additional ways to access content, improving overall comprehension, and supporting all audiences. Captions improve comprehension by removing common miscommunications stemming from variations in:
hearing ability
speaker accent
audio quality
listener’s primary language
complexity of the subject matter
missed audio content
Terminology
Captioning terminology can be confusing. Establishing a shared understanding of the terms used is important, for example:
Captions are synchronized text of spoken dialogue, sound effects, or other audio cues that appear within a video. Closed captions can be turned on and off, open captions are locked on.
Transcripts document the content from an audio recording (e.g., spoken dialogue, sound effects, or other audio cues) within a separate text file not time-synchronized to the video. Interactive transcripts are time-synchronized text versions of audio. They are typically searchable, and user selection of the text jumps the video to that specific location.
Captions and transcripts can be machine-generated (automated speech recognition or ASR), or generated by a human captionist, or both.
Enhanced automated speech recognition (ASR) captions use preprogrammed vocabulary and other preparatory materials to maintain an accuracy rate of at least 95%.
Live captions are generated in real time as audio is captured during a live event, while post-production captions are generated for previously recorded content.
Live Captioning
Live captioning transmits speech and audio into real-time text by generating text that corresponds to the auditory information. Live captions may come from a machine, a human, or a combination of both. When live captions are available (e.g., in Zoom, or Panopto), users can enable this feature to display all spoken content as live text.
Live captioning is required for any live, digital, public-facing event across the university (e.g., webinars, seminars, and town halls). Synchronous live captions can be integrated into Zoom or Microsoft Teams or provided via a separate browser window for the end user. As outlined below, Live Captioning is a requestable service for university faculty and staff.
Live Event Caption Planning
When planning a live event, it is important to plan for live captioning. Centralized funding covers the cost of CaptivateTM Automated Speech Recognition (ASR) captions for live events, at no cost to the event organizer. These captions are about 95% accurate without preparatory materials. Other options for live captions include Zoom or Microsoft Teams ASR captions. Please refer to the chart below to determine the best option for your live event.
| Functionality | Captivate ASR | Zoom | Microsoft Teams |
|---|---|---|---|
| Utilization | Schedule at least two business days in advance | Turn on at the start of the meeting | Turn on at the start of the meeting |
| Accuracy | Higher captioning accuracy
| Lower captioning accuracy
| Lower captioning accuracy
|
| Non-speech audio events | Included (e.g., door closing, dog barking, etc.) | Not included | Not included |
| Speaker identification | Included | Included | Included in default
|
| Appearance Settings | Livestream link allows the user to change font type, size, and color contrast
| Option to change font type, size, and color contrast
| Option to change font type, size, and color contrast
|
| Speed | Real-time, in-meeting | Real-time, in-meeting | Real-time, in-meeting |
| Cost | No cost to VT user | No cost to VT user | No cost to VT user |
| Post-event Transcript | Post-event transcript provided | Download the transcript prior to ending the Zoom meeting | Download the transcript from the meeting event calendar |
If using CaptivateTM captions, please submit a Live Captions Request form. When possible, event planners can increase captioning accuracy for their event by providing event preparation materials at least two business days in advance. Examples of preparation materials include:
Technical vocabulary
List of speaker names
Speaker biographies
Presentation materials (e.g., slide deck)
Event program or script
For support in determining the best live captioning plan for an event, please email captioning@vt.edu.
Post-Production Captioning
For pre-recorded videos, such as news announcements or course lectures, disability legislation requires post-production captions. For all videos uploaded to Panopto, machine captions are generated using automatic speech recognition (ASR) software. Accuracy is typically around 95%, though factors such as audio quality, the speaker, and specialized or complex terminology may impact accuracy. Any recordings shared through a public-facing website or Canvas course should be edited for accuracy using Panopto’s built-in captioning editor. Please refer to the Captioning in the Panopto 4Help article for more information.
Transcripts
Transcripts are necessary for audio-only content, such as podcasts. Audio and video files in Panopto include an interactive transcript that viewers can access by selecting “Captions” on the player’s menu bar. When using Live Captioning services integrated with Zoom or Microsoft Teams, a live transcript is also available. Users can order transcripts for audio recordings outside of Panopto by submitting the Transcription Request form.
Questions
Please email captioning@vt.edu with questions about captioning or transcripts.