Concept 1 / 1
Key Concepts to Memorize
Physiology of phonation:
- Vocal folds = primary sound generator; ventricular folds are passive
- Larynx serves two functions: (1) swallowing (3-step valve), (2) speech/singing
- Fundamental frequency of voice: f0 = 150–1500 Hz
- Physics of phonation = Fluid-Structure-Acoustic Interaction (FSAI)
Dysphonia = voice disorder:
- Symptoms: hoarseness, decreased load capacity, incomplete glottis closure, asymmetric oscillations
- Organic Dysphonia (structural cause): malformation, trauma, inflammation, malignant/benign growth (e.g., polyp, squamous cell carcinoma)
- Functional Dysphonia (no primary organ findings): over/incorrect loading, multiple combined causes
Clinical diagnostics of dysphonia (multimodal):
- 2D visualization: Laryngoscopy, Stroboscopy, Highspeed endoscopy
- Acoustic signal analysis: Voice field measurement, irregularity parameters
- Self + expert evaluation
- ElectroGlottoGraphy (EGG)
- Key limitation: in vivo examination of sound generation during phonation is NOT completely possible
Deep learning in laryngoscopy (3 tasks):
1. Localization of the glottis and vocal folds
2. Automatic segmentation of the glottis area
3. Classification of tissue type, organic disorder, etc.
- BAGLES benchmark: 7 hospitals (EU + US), 640 records, 5 cameras, 59,250 images with segmentation
3 types of larynx models:
| Model Type | Degree of Reality | AI Support | Data Density |
|---|---|---|---|
| Ex vivo | Highest | Some | Medium |
| Synthetic (silicone) | Medium | Medium | Medium |
| Computational (CFD) | Lowest | Highest | Highest |
AI-supported CFD simulations:
- Classical CFD = extremely slow: 140 cores, 10h per cycle → 100h for 10 cycles
- Solution: SIREN (Implicit Neural Representations with periodic activation functions)
- SIREN enables: (1) increase spatial resolution, (2) increase temporal resolution, (3) future prediction of flow fields
Take-home message: AI in biomedical science goes far beyond MRI/CT postprocessing.