Nimendra's Notes ðŠī
Search
Search
Dark mode
Light mode
Reader mode
Explorer
Home
âŊ
00.Fleeting Notes
âŊ
AI-LLM
âŊ
AI Agents
âŊ
2. LLMs and AI Agents
2. LLMs and AI Agents
Table of Contents
What You Should Know About Large Language Models
LLMs inside agents
Input and Output Tokens
From Language Modeling to Powering Agents
System Prompt
Multi-turn Conversations
Tool Use
Thought â Action â Observation loop
Training a Large Language Model
Pre-training â Language Modeling
Post-training
Supervised Fine-Tuning
Mental model
The one thing to remember
Reinforcement Learning
Group Relative Policy Optimization â GRPO
Transformer Architecture
Tokenizer
Transformer blocks
LM Head
Decoding and Temperature
Higher temperature
Processing Through Transformer Blocks
Context Length
KV / Prompt Caching
Inside a Transformer Block
Feed-Forward Neural Network
Dense Models vs Mixture-of-Experts
Self-Attention â High-Level Intuition
A Deeper Dive Into Large Language Models
How Self-Attention Works
Relevance Scoring
Combining Information
KV-Caching Revisited
KV cache trade-off
More Efficient Self-Attention
Multi-head Latent Attention â MLA
MLA and positional information
Simplified MLA intuition
34. MHA vs GQA vs MQA vs MLA
MHA â Multi-Head Attention
GQA â Grouped-Query Attention
MQA â Multi-Query Attention
MLA â Multi-head Latent Attention
35. DeepSeek Sparse Attention â DSA
Lightning Indexer
Top-K Selector
Main result
36. Mixture of Experts â MoE
Experts
Router
37. Expert Load Balancing
Expert Capacity
Noise and Auxiliary Loss
38. Sparse Parameters vs Active Parameters
39. DeepSeek-R1 MoE Example
40. Example MoE Models
41. Whole-Chapter Mental Model
42. Agent Developer Mental Model
43. Key Distinctions to Remember
Graph View