Tag Archives | attention

Inferencing the Transformer Model

By Stefania Cristina on January 6, 2023 in Attention 11

We have seen how to train the Transformer model on a dataset of English and German sentence pairs and how to plot the training and validation loss curves to diagnose the model’s learning performance and decide at which epoch to run inference on the trained model. We are now ready to run inference on the […]

Plotting the Training and Validation Loss Curves for the Transformer Model

By Stefania Cristina on January 6, 2023 in Attention 7

We have previously seen how to train the Transformer model for neural machine translation. Before moving on to inferencing the trained model, let us first explore how to modify the training code slightly to be able to plot the training and validation loss curves that can be generated during the learning process. The training and […]

Training the Transformer Model

By Stefania Cristina on January 6, 2023 in Attention 44

We have put together the complete Transformer model, and now we are ready to train it for neural machine translation. We shall use a training dataset for this purpose, which contains short English and German sentence pairs. We will also revisit the role of masking in computing the accuracy and loss metrics during the training […]

Joining the Transformer Encoder and Decoder Plus Masking

By Stefania Cristina on January 6, 2023 in Attention 32

We have arrived at a point where we have implemented and tested the Transformer encoder and decoder separately, and we may now join the two together into a complete model. We will also see how to create padding and look-ahead masks by which we will suppress the input values that will not be considered in […]

Implementing the Transformer Decoder from Scratch in TensorFlow and Keras

By Stefania Cristina on January 6, 2023 in Attention 11

There are many similarities between the Transformer encoder and decoder, such as their implementation of multi-head attention, layer normalization, and a fully connected feed-forward network as their final sub-layer. Having implemented the Transformer encoder, we will now go ahead and apply our knowledge in implementing the Transformer decoder as a further step toward implementing the […]

Implementing the Transformer Encoder from Scratch in TensorFlow and Keras

By Stefania Cristina on January 6, 2023 in Attention 5

Having seen how to implement the scaled dot-product attention and integrate it within the multi-head attention of the Transformer model, let’s progress one step further toward implementing a complete Transformer model by applying its encoder. Our end goal remains to apply the complete model to Natural Language Processing (NLP). In this tutorial, you will discover how […]

The Vision Transformer Model

By Stefania Cristina on January 6, 2023 in Attention 5

With the Transformer architecture revolutionizing the implementation of attention, and achieving very promising results in the natural language processing domain, it was only a matter of time before we could see its application in the computer vision domain too. This was eventually achieved with the implementation of the Vision Transformer (ViT). In this tutorial, you […]

How to Implement Multi-Head Attention from Scratch in TensorFlow and Keras

By Stefania Cristina on January 6, 2023 in Attention 27

We have already familiarized ourselves with the theory behind the Transformer model and its attention mechanism. We have already started our journey of implementing a complete model by seeing how to implement the scaled-dot product attention. We shall now progress one step further into our journey by encapsulating the scaled-dot product attention into a multi-head […]

How to Implement Scaled Dot-Product Attention from Scratch in TensorFlow and Keras

By Stefania Cristina on January 6, 2023 in Attention 5

Having familiarized ourselves with the theory behind the Transformer model and its attention mechanism, we’ll start our journey of implementing a complete Transformer model by first seeing how to implement the scaled-dot product attention. The scaled dot-product attention is an integral part of the multi-head attention, which, in turn, is an important component of both […]

muhammad-murtaza-ghani-CIVbJZR8aAk-unsplash

A Gentle Introduction to Positional Encoding in Transformer Models, Part 1

By Mehreen Saeed on January 6, 2023 in Attention 43

In languages, the order of the words and their position in a sentence really matters. The meaning of the entire sentence can change if the words are re-ordered. When implementing NLP solutions, recurrent neural networks have an inbuilt mechanism that deals with the order of sequences. The transformer model, however, does not use recurrence or […]

1 2 Next →