The speech group is a sub-group of the Center for Processing Speech and Images (PSI). The ESAT speech group is internationally recognized as a leading research center in large vocabulary speech recognition, spoken vocabulary acquisition, noise robustness and exemplar-based recognition.
The group also maintains and develops a fully in-house developed state-of-the-art continuous speech, large vocabulary, speaker independent recognition system.
Project
Your research will be conducted within the CAMETRON project which aims at building a virtual director for audiovisual recordings involving several cameras and microphones. Active steering of the recording nodes starts from an audiovisual scene analysis and understanding. Your research domain is audio scene analysis as well a audiovisual integration. Audio processing involves three related problems: locating the source, segragating its signal from other sound sources and identifying it, for instance as a speaker that was observed earlier in the audiovisual production.
Knowledge of the solution of one problem as well as knowledge from the visual modality helps solving the other problems. We therefore will study a framework to solve these problems jointly using a factorization method. Spectro-temporal observations of the audio mixtures observed in each recording nodes are written as linear combinations of typical spectro-temporal patterns that are typical for a particular source. These spectro-temporal patterns are additionally subject to modifications induced by the acoustic environment (direction of arrival, reverberation, noise, ...). The project involves plugging in the right mathematical models in this framework and estimating its components from data.
Junior researcher in speech processing
label
Burse
calendar_month
2013-07-03, 00:00
autorenew
2025-09-29, 17:01
history_edu
Diana Ignat