llm ai agents reading/researchpaper

Chain of Thought (CoT)
3. Agentic AI and Autonomous Systems

The Idea at a Glance

flowchart LR
    Q["Question / Task"] --> T1["Thought: what do I need next?"]
    T1 --> A1["Action: call a tool (search / act)"]
    A1 --> O1["Observation: read the result"]
    O1 --> T2["Thought: updated plan"]
    T2 --> A2["Action"]
    A2 --> O2["Observation"]
    O2 -->|"loop until solved"| T3["Thought"]
    O2 -.->|"enough info"| ANS["Final Answer"]

The loop keeps repeating: reason, act, observe, reason again — until the agent has enough information to answer.

Why Both? (CoT vs Act vs ReAct)

ApproachReasoningActingProblem
CoT (think only)✅❌Hallucinates facts; no access to real information
Act-only (act only)❌✅No planning; repeats wrong actions
ReAct (think + act)âś…âś…Grounded in real observations, plans as it goes

The key trick: the model’s thoughts and actions are interleaved in one trajectory, so its reasoning is guided by what it actually observes.

Key Pieces

  • Thought — free-form reasoning: decompose goals, extract facts, adjust plans
  • Action — calling an external tool (e.g. search query, environment step)
  • Observation — the tool’s result, fed back into the next thought
  • No fine-tuning needed — works by few-shot prompting a frozen LLM (PaLM-540B) with 1–6 human-written example trajectories

Results (from the paper)

  • More grounded, fewer hallucinations — false-positive rate 6% vs 14% for CoT on HotpotQA; hallucination is CoT’s main failure mode (56%)
  • Beats Act-only on both knowledge QA and decision-making (ALFWorld) — reasoning guides better actions
  • vs CoT: better on FEVER fact verification (60.9 vs 56.3), slightly behind on HotpotQA (27.4 vs 29.4) — but with far fewer made-up facts
  • Weakness: repetitive thought/action loops, and it depends on search returning good information

Why It Matters

ReAct is the foundation of modern tool-using agents (the industry-standard pattern for multi-tool workflows), connecting tool interfaces to an iterative agent loop. Any agent that “thinks, calls a tool, and reacts to the result” is following the ReAct loop.