Share it

Over the last decade, Deep Learning has swept through almost every area of ICT, penetrating so quickly and pervasively into business and social domains that it has enabled unprecedented new capabilities. Some of them are now as commonplace as the controversial photo-realistic fake portraits, also known as deepfakes.

Here, Generative Neural Networks have undergone intensive and constant improvement in their capabilities, including several paradigm shifts. Some of their variants allow impressive techniques for Multimodal Domain Translation (MDT), whereby for example an image can be generated from a simple textual description or a soundless video excerpt can be populated with a synthetic and plausible audio track. In the state of the art, Conditional Generative Adversarial Neural Networks (cGAN) have some benefits and demonstrate excellent results compared to Variational Autoencoders (VAE) in a competitive race where Transformers have recently burst onto the scene with high expectations.

In this masterclass, we will go over several creative MDT applications and examples of image-to-image, text-to-sound and video-to-sound. We will also discuss several techniques to improve and stabilise the training of cGANs such as conditioning augmentation, spectral normalisation, adaptive data augmentation, gradient penalties, self-attention, class conditional normalisation and conditional projection classifiers. A final look at attentional architectures, which are the core of the Transformers paradigm, will showcase their unique capabilities as translators over and above Natural Language Processing such as GPT-3 and as generators in general.

Addressed to:

– Computer vision experts – Professionals in multimedia content creation industries (image, audio, music) – Professionals in the design, advertising and creative sectors, digital artists – Audiobook industry – ICT industry professionals in general

Programa

1. Creative applications and some examples
2. Techniques to improve and stabilise cGAN training
3. Attentional architectures
4. Q&A

Masterclass taught by Eurecat (CIDAI core partner)

Taught by:

Rafael Redondo
Senior Researcher on Computer Vision at Eurecat
CIDAI