A901: Identify features and capabilities of speech recognition and speech synthesis

1. Speech Recognition (Speech-to-Text)
📌 What it does

Converts spoken audio into written text

🔑 Key Features & Capabilities
Transcribes conversations in real time or from recordings
Supports multiple languages
Can identify punctuation (in advanced systems)
Speaker diarization (who said what)
Keyword recognition
âś… Example scenarios
Transcribing call centre recordings
Voice commands (“Open dashboard”)
Meeting transcription

👉 Exam clue:

“Convert speech/audio into text”
➡️ Speech recognition

🔊 2. Speech Synthesis (Text-to-Speech)
📌 What it does

Converts written text into spoken audio

🔑 Key Features & Capabilities
Natural-sounding voices
Multiple voice styles (male/female, tone, emotion)
Adjustable speed, pitch, and pronunciation
Multilingual support
âś… Example scenarios
Reading out notifications or alerts
Accessibility (screen readers)
Voice assistants

👉 Exam clue:

“Generate spoken audio from text”
➡️ Speech synthesis

⚖️ 3. Key Differences (Very Important for Exam)
Capability Speech Recognition Speech Synthesis
Input Audio Text
Output Text Audio
Purpose Understand speech Generate speech

đź§© 4. Combined Use (Often Tested)

Many real-world solutions use both together:

🎯 Example

A voice assistant that listens to a user and responds verbally

User speaks → Speech recognition
System responds → Speech synthesis

⚠️ 5. Common Exam Traps
Confusing direction:
Speech → Text = Recognition
Text → Speech = Synthesis
Thinking speech AI is only transcription
❌ It also includes voice generation

đź§  6. Simple Memory Trick

“Hear vs Speak”

Hear → Recognition (understand speech)
Speak → Synthesis (produce speech)
âś… Summary

To answer AI-901 questions in this area:

Identify input and output type
Map to:
Audio → Text → Speech Recognition
Text → Audio → Speech Synthesis
Recognise key capabilities like:
Transcription, translation → Recognition
Voice generation, playback → Synthesis