Abstract
FNN (Feed-Forward Neural Network) and RNN (Recurrent Neural Network) are the two classic neural network families. FNN processes each input independently β data flows one way, no memory. RNN has loops: it keeps a hidden state that carries memory of what came before, so it can handle sequences where order matters. Every modern LLM (Transformer) descends from these ideas β and both are trained with Backpropagation.

Feed-Forward Neural Network (FNN)
Definition
- The simplest type of neural network.
- Data flows in one direction only: Input β Hidden Layer(s) β Output.
- There are no feedback loops or memory.
flowchart LR I["Input"] --> H1["Hidden 1"] --> H2["Hidden 2"] --> O["Output"] style I fill:#d4e6f1 style O fill:#d5f5e3
Characteristics
- No memory of previous inputs β each sample processed independently.
- Best for static (independent) data.
- Simple architecture; faster and easier to train.
Applications
- Image classification
- Object detection
- Handwritten digit recognition (MNIST)
- Medical diagnosis
- Credit scoring
Recurrent Neural Network (RNN)
Definition
- A neural network designed for sequential data.
- Contains feedback connections (recurrent loops) that allow it to remember previous inputs.
flowchart LR subgraph one["Time step 1"] X1["xβ"] --> R1["Hidden hβ"] end subgraph two["Time step 2"] X2["xβ"] --> R2["Hidden hβ"] end subgraph three["Time step 3"] X3["xβ"] --> R3["Hidden hβ"] end R1 -->|"memory"| R2 R2 -->|"memory"| R3
Characteristics
- Has memory using hidden states β the same network is reused at each time step.
- Processes data step by step; order and context matter.
- More complex and slower to train.
- Can suffer from vanishing/exploding gradient problems (see Backpropagation Β§9).
Applications
- Language translation
- Speech recognition
- Sentiment analysis
- Time-series forecasting
- Stock price prediction
- Weather forecasting
Key Differences
| Feature | FNN | RNN |
|---|---|---|
| Data Flow | One-way (Input β Output) | Recurrent/feedback loops |
| Memory | No memory | Yes (hidden state) |
| Data Type | Static, independent | Sequential / time-series |
| Context Awareness | No | Yes |
| Training Speed | Faster | Slower |
| Complexity | Simple | More complex |
| Common Problems | None significant | Vanishing/exploding gradients |
When to Use
Use FNN when:
- Data samples are independent.
- No previous information is needed.
- Example: Image classification or customer credit scoring.
Use RNN when:
- Data is sequential.
- Previous information affects future predictions.
- Example: Predicting the next word in a sentence or forecasting stock prices.
Simple Example
FNN
Input: Image of a cat β Predicts: Cat
Each image is processed independently.
RNN
Input Sentence:
βI love learning ___β
The model remembers the previous words βI love learningβ and predicts: β AI
The prediction depends on the previous context.
Related Notes
- Backpropagation β the training algorithm behind both architectures
- Mixture of Experts (MoE) β how modern models scale the FNN feedforward layer
- Chain of Thought (CoT) β reasoning techniques for todayβs LLMs
- RLHF - Reinforcement Learning from Human Feedback β how LLMs are aligned after training
- KV cache β the attention state caching that makes transformer (and hence LLM) decoding efficient