Transformer من الصفر
أوت 2025 – سبتمبر 2025
تنفيذ بنية Transformer بـ PyTorch، بما يشمل التضمينات، والانتباه متعدد الرؤوس، والترميز الموضعي، وكتل المُرمِّز والمفكِّك، اتباعًا لـ «Attention Is All You Need».
I wanted to understand the Transformer from the inside rather than through a library call, so I implemented the architecture of "Attention Is All You Need" from scratch in PyTorch, working alone.
What I built
I followed Umar Jamil's walkthrough, Coding a Transformer from scratch on PyTorch , alongside the paper, and adapted it to English-to-French translation on the OPUS Books dataset. The code covers the embeddings, positional encoding, multi-head attention and the encoder and decoder blocks, then word-level tokenizers, the training and validation loops, and greedy decoding. I never finished training: without a GPU, a full run was out of reach.
What I took from it
I came out understanding attention in depth, and the distance between the paper and working code. Many formulas cannot be copied as written. The positional-encoding denominator is computed through exponentials and logarithms to stay numerically stable, and the attention heads are views of one tensor instead of separate matrices, because memory sets the limits that the equations ignore.