ai llm machinelearning agents

The Idea at a Glance

flowchart LR
    T["Token"] --> R["Router / Gate (tiny learned network)"]
    R -->|"scores all experts"| E1["Expert 1"]
    R -->|"top-k only"| E2["Expert 2"]
    R -.->|"skipped"| E3["Expert 3"]
    E1 --> O["Weighted output"]
    E2 --> O
    style E3 stroke-dasharray: 5 5

8 experts exist → a token activates only 2. The rest stay idle.

Why Bother?

Dense modelMoE model
Total parametersAll used per tokenMany, mostly idle
Active per tokenAllOnly top-k
Compute per tokenHighRelatively low
Capability—Close to a model with all parameters

MoE = large total capacity (what it knows) at low per-token cost (how much it computes).

Key Pieces

  • Expert — an FFN, structurally the same as the dense FFN it replaces. Specialization is learned, not hand-assigned.
  • Router / gate — a small learned network that scores every expert per token (top-k executed, rest skipped → sparse).
  • Load balancing — a loss that stops a few experts from hogging all tokens.
  • Two routing styles — tokens choose experts (standard) vs experts choose tokens.
  • Upcycling — building MoE from an existing model by copying its FFN into multiple experts.

Real Examples

  • Mixtral 8x7B — 8 experts, top-2 routing
  • DeepSeek-V3 — 671B total, only ~37B active per token
  • Qwen3-235B-A22B — 235B total, 22B activated (read the name: total-Aactive)

Pros & Cons

Pros: more capacity for roughly dense-compute cost; fewer FLOPs per token; scaling without proportional cost.

Cons: more memory (all experts must fit); router + load-balancing complexity; communication overhead when experts are sharded across machines; harder to serve than dense models, so inference optimization matters.


Deep Dive