Nimendra's Notes 🪴

Home

❯

00.Fleeting Notes

❯

AI-LLM

❯

Hands-on Large Language Models

❯

Attention is All You Need (Transformers)

Attention is All You Need (Transformers)


Backlinks

  • 2. LLMs and AI Agents
  • 1. Introduction to LLMs
  • 4. Architectures (LLM)
  • 9. Inference Optimization
  • 1. An Introduction to LLM
  • Decoder-Only Transformers (GPT)
  • Hands-on Large Language Models
  • KV cache
  • Backpropagation
  • FNN and RNN
  • How Compaction Works in Pi Agent
  • Introduction
  • Why was the Transformer Introduced ?
  • Problems with RNNs/LSTMs
  • The Solution: Self-Attention
  • Transformer
  • What is the Transformer ?
  • Architecture
  • Encoder
  • Self Attention
  • Decoder
  • Why the Transformer is Better than RNNs
  • Softmax Function
  • The Annotated Transformer
  • Transformer Architecture: A Socratic Guide
  • 1. Why do we need Transformers?
  • The Transformer Idea
  • 2. What is Self-Attention?
  • 3. From Words to Numbers
  • 4. Why Matrix Multiplication?
  • 5. How are Attention Scores Created?
  • 6. Softmax Converts Scores into Attention
  • Softmax Function
  • 7. What Happens to Value (V)?
  • Complete Self-Attention Pipeline
  • 8. How Does the Transformer Know?
  • 9. Residual Connection
  • 10. Feed Forward Network
  • 11. Why Many Transformer Layers?
  • Question
  • 12. Masked Self-Attention
  • How Masking Works
  • Normal vs Masked Self-Attention

Graph View

Created with Quartz v5.0.0 © 2026

  • Blog
  • GitHub
  • X(Twitter)