Automatic speech recognition turns what is said into text, almost in real time. Video calls, online videos and face-to-face conversations can be captioned without any preparation.
For deaf and hard-of-hearing people, it gives immediate access to exchanges that used to be hard to follow. It also helps in a noisy place or in a language one does not master.
Automatic transcription still makes mistakes: proper nouns, technical vocabulary, accents and overlapping speech remain tricky. For prepared content, captions reviewed by a person are still the reference.
On the design side, plan where captions sit, sufficient contrast, an adjustable size and a way to read back what has just been said.