{"id":115,"date":"2022-07-21T19:23:53","date_gmt":"2022-07-21T18:23:53","guid":{"rendered":"https:\/\/wp.coventry.domains\/e2edu\/?page_id=115"},"modified":"2022-08-23T18:14:03","modified_gmt":"2022-08-23T17:14:03","slug":"autoencoder","status":"publish","type":"page","link":"https:\/\/wp.coventry.domains\/e2edu\/autoencoder\/","title":{"rendered":"Autoencoder"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">The following article introduces Autoencoders that operate on images. Nevertheless, many aspects described here also apply for Autoencoders that operate different types of data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Autoencoders are Machine-Learning models that learn to replicate data that is provided as input in the output the generate. Accordingly, Autoencoders operate as copy machines.  What makes an Autoencoder&#8217;s task both challenging and interesting is the fact that it has to deal with an information bottleneck. Data that passes through this bottleneck is of much lower dimension than that of the input and output data. In order to succeed in their data replication task, Autoencoders have to learn an effective abstraction of the data. This abstracting plays the role of an encoding of the data from which it can be faithfully reconstructed. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"868\" height=\"601\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_bottleneck.png\" alt=\"\" class=\"wp-image-747\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_bottleneck.png 868w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_bottleneck-300x208.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_bottleneck-768x532.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_bottleneck-788x546.png 788w\" sizes=\"auto, (max-width: 868px) 100vw, 868px\" \/><figcaption>Bottleneck Layer in an Autoencoder. \u00a9Hmrishav Bandyopadhyay<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">Autoencoders are the somewhat related to Generative Adversarial Networks. They are also frequently used for creating synthetic images. Autoencoders consist of two models, named Encoder and Decoder. The Decoder plays the same role as a Generator in a GAN. The Decoder takes as input a low dimensional vector and generates a synthetic image from it. Contrary to a GAN, this vector does not contain random values. Instead, it is generated by the Encoder. The Encoder takes an image from a dataset and generates a latent vector that represents the image&#8217;s encoding. This vector is then passed as input to the Decoder which generates a synthetic replica of the original image. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"242\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder-1024x242.png\" alt=\"\" class=\"wp-image-753\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder-1024x242.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder-300x71.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder-768x182.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder-788x186.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/AutoEncoder_Encoder_Decoder.png 1400w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Data Processing Pipeline of an Autoencoder that Operates on Images. From left to right: an original image is passed as input into the Encoder with generates an encoding of the image. The encoding is passed through the decoder which generates a synthetic image as output.  \u00a9Hmrishav Bandyopadhyay<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">The Encoder and Decoder parts of an Autoencoder are implemented as neural networks. In the case of Autoencoders that operate on images, these networks are convolutional neural networks. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"850\" height=\"311\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Convolution_Deconvolution.png\" alt=\"\" class=\"wp-image-772\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Convolution_Deconvolution.png 850w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Convolution_Deconvolution-300x110.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Convolution_Deconvolution-768x281.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Convolution_Deconvolution-788x288.png 788w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><figcaption>Convolution and Deconvolution Layers in an Autoencoder. The left side depicts the Encoder, the right side depicts the Decoder. \u00a9Vaibhav Kumar<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">During training, the Autoencoder tries  to minimise the reconstruction error between the original and synthetic image. Both the Encoder and Decoder models are trained in parallel. After training, the Encoder is usually no longer used. Instead, latent vectors are created either randomly or using vector arithmetic are then input into the Decoder to generate new synthetic images. <\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Latent Vector Arithmetic<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">One of the fascinating aspects of working with machine learning models that operate on latent codes such as GANs and Autoencoders is to explore how variations of latent codes that encode a known image might lead to the generation of interesting new synthetic images. <\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Vector Interpolation<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">One approach of creating such variations is to interpolate between two or more known encodings for images. By interpolating, new encodings are created that, once decoded, lead to synthetic images that combine some of the characteristics of the original images. This combination is not the same as a simple blending between images (which is obtained when the images themselves and not their encodings are interpolated). Instead, it is the codified abstractions used in the encodings that are interpolated and that manifest as a blending of higher level abstractions of images. This includes for example a morphing of the shape or type of the depicted objects. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"512\" height=\"512\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Bilinear-interpolation-on-latent-space.png\" alt=\"\" class=\"wp-image-787\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Bilinear-interpolation-on-latent-space.png 512w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Bilinear-interpolation-on-latent-space-300x300.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Autoencoder_Bilinear-interpolation-on-latent-space-150x150.png 150w\" sizes=\"auto, (max-width: 512px) 100vw, 512px\" \/><figcaption>Bilinear Interpolation of Encodings of Images of Faces. \u00a9Mathijs Pieters and Marco A. Wiering, 2018<\/figcaption><\/figure>\n<\/div>\n\n\n<h3 class=\"wp-block-heading\">Vector Subtraction and Addition<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Another approach is to conduct vector operations on encodings, such as adding or removing differences between two encodings to another encoding. Following this approach, specific properties of an image can be selectively changed. The typical method for doing this is as follows: The encodings of all the images that contain one version of the selected property  (such as a smiling facial expression) are combined into one averaged encoding. The same is done for the other version of the selected property (such as a non-smiling facial expression). Then a difference encoding is calculated by subtracting the average encoding for non-smiling faces from the average encoding from smiling faces. This difference encoding, when added to any encoding of a non-smiling face, will create, once decoded, the same face but with a smiling facial expression. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"531\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN-1024x531.png\" alt=\"\" class=\"wp-image-790\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN-1024x531.png 1024w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN-300x155.png 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN-768x398.png 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN-788x408.png 788w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/Example-of-Vector-Arithmetic-on-Points-in-the-Latent-Space-for-Generating-Faces-with-a-GAN.png 1104w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Conducting Vector Operations with Encodings of Images of Faces. \u00a9Jason Brownlee<\/figcaption><\/figure>\n<\/div>\n\n\n<h2 class=\"wp-block-heading\">Latent Space Organisation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Unfortunately, a vanilla Autoencoder such as the one described above won&#8217;t create encodings that are suitable for latent vector arithmetic. Variations of latent vectors can only be decoded into meaningful images, if the distribution of latent vectors within latent space fulfils two criteria. First, the latent space doesn&#8217;t contain any gaps. Gaps in this case are regions in latent space that don&#8217;t encode meaningful images. Second, images that are similar to each other should have their encodings lie close to each other in latent space. Then the Euclidean distance between encodings is a measure of similarity between the images they encode. If this is not the case, then a gradual interpolation between known encodings of images won&#8217;t lead to images that gradually increase and decrease the features of the original images. Instead, these features would arbitrarily change.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"642\" height=\"611\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/structuring_latent_space.jpg\" alt=\"\" class=\"wp-image-798\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/structuring_latent_space.jpg 642w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/structuring_latent_space-300x286.jpg 300w\" sizes=\"auto, (max-width: 642px) 100vw, 642px\" \/><figcaption>Latent Space Distributions for Different Types of Autoencoders. AE stands for vanilla Autoencoder, VAE for Variational Autoencoder, AAE for Adversarial Autoencoder, and SAE for Structuring Autoencoder.  \u00a9Marco Rudolph et. al. 2019<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">Two types of Autoencoders that have been devised to improve the organisation of latent space are briefly introduced in the remainder of this article. <\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Variational Autoencoder<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Variational Autoencoders (VAE) differ from vanilla Autoencoders (AE) in that the former are probabilistic and the latter deterministic. In a AE, the Encoder generates from an input image a single deterministic encoding. This encoding is directly decoded by the Decoder into a synthetic image. In a VAEs, the Encoder generates from an input image the parameters for a gaussian distribution of encodings. These parameters are the mean and the standard deviation for each of the dimensions of the encoding. The Decoder then samples from this distribution to obtain an encoding that is then decoded into a synthetic image. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The error that a VAE tries to minimise is a combination of two errors, a reconstruction error (which is the same as used when training a AE) and a <a rel=\"noreferrer noopener\" href=\"https:\/\/en.wikipedia.org\/wiki\/Kullback%E2%80%93Leibler_divergence\" target=\"_blank\">Kullback-Leibler divergence<\/a>. In a nutshell, this divergence is large if two probability distributions don&#8217;t match, otherwise it is small. <\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"720\" height=\"444\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/KL_Divergence_Matching_Distributions.png\" alt=\"\" class=\"wp-image-808\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/KL_Divergence_Matching_Distributions.png 720w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/KL_Divergence_Matching_Distributions-300x185.png 300w\" sizes=\"auto, (max-width: 720px) 100vw, 720px\" \/><figcaption>Kullback-Leibler Divergence between Two Example Distributions p(x) and q(x). \u00a9datumorphism.leima.is<\/figcaption><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">In the case of a VAE, the two distributions that should match are the distribution generated by the Encoder (the predicted distribution) and a set of Gaussian distributions whose mean is at zero and that have a standard deviation of one. Minimising the divergence causes the distributions of the encodings created by the Encoder to cluster close together while the Euclidean distance between the encodings is representative of the similarity between the encoded images. For a more in depth and mathematical explanation of VAE, the article &#8220;<a rel=\"noreferrer noopener\" href=\"https:\/\/datahacker.rs\/gans-004-variational-autoencoders-in-depth-explained\/\" target=\"_blank\">Variational Autoencoders &#8211; in depth explained<\/a>&#8221; is a good resource. <\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Adversarial Autoencoder<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">An alternative approach to obtain a well organised latent space is to add an Discriminator model to the Autoencoder.  The role of the Discriminator is similar to the role of the Critique in a Generative Adversarial Network (GAN), that is to distinguish between real and fake instances of data. For this reason, this type of Autoencoders is named Adversarial Autoencoder (AAE). In an AAE, the Discriminator&#8217;s task is to distinguish between values taken from a real Gaussian Distribution and values generated by the Encoder. Accordingly, the Discriminator is a classifier for real versus fake random variables. <\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"850\" height=\"633\" src=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/An-Adversarial-autoencoder-network-15_W640.jpg\" alt=\"\" class=\"wp-image-828\" srcset=\"https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/An-Adversarial-autoencoder-network-15_W640.jpg 850w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/An-Adversarial-autoencoder-network-15_W640-300x223.jpg 300w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/An-Adversarial-autoencoder-network-15_W640-768x572.jpg 768w, https:\/\/wp.coventry.domains\/e2edu\/wp-content\/uploads\/sites\/3486\/2022\/08\/An-Adversarial-autoencoder-network-15_W640-788x587.jpg 788w\" sizes=\"auto, (max-width: 850px) 100vw, 850px\" \/><figcaption>Encoder, Decoder, and Discriminator in an Adversarial Autoencoder. \u00a9Vahid Mirjalili<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For a more in depth and mathematical explanation of AAE, the article &#8220;<a href=\"https:\/\/medium.com\/vitrox-publication\/adversarial-auto-encoder-aae-a3fc86f71758\" target=\"_blank\" rel=\"noreferrer noopener\">Adversarial Auto Encoder (AAE)<\/a>&#8221; is a good resource. <\/p>\n","protected":false},"excerpt":{"rendered":"<p>The following article introduces Autoencoders that operate on images. Nevertheless, many aspects described here also apply for Autoencoders that operate different types of data. Autoencoders are Machine-Learning models that learn to replicate data that is provided as input in the output the generate. Accordingly, Autoencoders operate as copy machines. What makes an Autoencoder&#8217;s task both [&hellip;]<\/p>\n","protected":false},"author":2154,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"_coblocks_attr":"","_coblocks_dimensions":"","_coblocks_responsive_height":"","_coblocks_accordion_ie_support":"","footnotes":""},"class_list":["post-115","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/115","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/users\/2154"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/comments?post=115"}],"version-history":[{"count":74,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/115\/revisions"}],"predecessor-version":[{"id":3074,"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/pages\/115\/revisions\/3074"}],"wp:attachment":[{"href":"https:\/\/wp.coventry.domains\/e2edu\/wp-json\/wp\/v2\/media?parent=115"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}