Nimendra's Notes ðŸŠī

Home

âŊ

00.Fleeting Notes

âŊ

AI-LLM

âŊ

AI Agents

âŊ

2. LLMs and AI Agents

2. LLMs and AI Agents


Backlinks

  • 0. An Illustrated Guide to AI Agents
  • What You Should Know About Large Language Models
  • LLMs inside agents
  • Input and Output Tokens
  • From Language Modeling to Powering Agents
  • System Prompt
  • Multi-turn Conversations
  • Tool Use
  • Thought → Action → Observation loop
  • Training a Large Language Model
  • Pre-training — Language Modeling
  • Post-training
  • Supervised Fine-Tuning
  • Mental model
  • The one thing to remember
  • Reinforcement Learning
  • Group Relative Policy Optimization — GRPO
  • Transformer Architecture
  • Tokenizer
  • Transformer blocks
  • LM Head
  • Decoding and Temperature
  • Higher temperature
  • Processing Through Transformer Blocks
  • Context Length
  • KV / Prompt Caching
  • Inside a Transformer Block
  • Feed-Forward Neural Network
  • Dense Models vs Mixture-of-Experts
  • Self-Attention — High-Level Intuition
  • A Deeper Dive Into Large Language Models
  • How Self-Attention Works
  • Relevance Scoring
  • Combining Information
  • KV-Caching Revisited
  • KV cache trade-off
  • More Efficient Self-Attention
  • Multi-head Latent Attention — MLA
  • MLA and positional information
  • Simplified MLA intuition
  • 34. MHA vs GQA vs MQA vs MLA
  • MHA — Multi-Head Attention
  • GQA — Grouped-Query Attention
  • MQA — Multi-Query Attention
  • MLA — Multi-head Latent Attention
  • 35. DeepSeek Sparse Attention — DSA
  • Lightning Indexer
  • Top-K Selector
  • Main result
  • 36. Mixture of Experts — MoE
  • Experts
  • Router
  • 37. Expert Load Balancing
  • Expert Capacity
  • Noise and Auxiliary Loss
  • 38. Sparse Parameters vs Active Parameters
  • 39. DeepSeek-R1 MoE Example
  • 40. Example MoE Models
  • 41. Whole-Chapter Mental Model
  • 42. Agent Developer Mental Model
  • 43. Key Distinctions to Remember

Graph View

Created with Quartz v5.0.0 ÂĐ 2026

  • Blog
  • GitHub
  • X(Twitter)