Easy AI Tutorial: LLaMA Model
> An open-source large language model developed by Meta
This post is part of the Easy AI tutorial series on zhichai.net. It teaches the LLaMA model through interactive visualizations covering architecture, data flow, version evolution, and parameter comparison.
Architecture
The visualization lets you interactively click through each component:
- Tokenizer — Converts input text into token sequences using efficient subword encoding. LLaMA-3 uses a 128K-vocabulary tokenizer, supporting multilingual text and reducing sequence length.
- Embedding layer — Maps token IDs into high-dimensional vector space, including positional encoding information, to provide semantic representations.
- Decoder — The core component: a stack of Decoder Blocks in a Decoder-Only architecture with multi-head self-attention (masked to ensure causality) and feed-forward layers.
- Self-Attention — Computes relationships between positions via Query, Key, and Value vectors and softmax attention weights. LLaMA-2/3 introduce Grouped Query Attention (GQA) for long-sequence modeling.
- MLP (Feed-Forward Network) — Two fully connected layers with activation functions, plus residual connections, for nonlinear feature transformation.
- Output layer — A linear layer projecting to vocabulary size, producing a probability distribution over the next token, supporting multiple decoding strategies.
- Parameter scale: 7B, 8B, 13B, 30B, 65B, 70B, up to 400B parameters
- Training data growth: from 1T to over 15T tokens, providing a solid foundation for performance gains
- Context length: from the original limit to 4,096 and 8K, enabling longer documents and complex conversations
Data Flow Animation
Using the sample input, the animation shows the processing pipeline:
1. Text input — the raw string 2. Tokenization — splitting into token sequences 3. Embedding — conversion to high-dimensional vectors 4. Self-attention — computing word-to-word relevance 5. Feed-forward network — nonlinear feature refinement 6. Output generation — probability distribution predicting the next word
Evolution Timeline
| Version | Date | Key highlights | |---|---|---| | LLaMA-1 | Feb 2023 | First open large-scale LLM; pretrained on 1T tokens; Transformer Decoder-Only, RMSNorm, SwiGLU activation, rotary position embeddings; quickly popular in the open-source community | | LLaMA-2 | Jul 2023 | Pretraining data expanded to 2T tokens; context doubled to 4,096; Grouped Query Attention (GQA); improved safety alignment; commercially friendly open-source license | | LLaMA-3 | Apr 2024 | 8K long-context support; efficient 128K-vocabulary tokenizer; over 15T tokens of training data; enhanced multilingual and instruction-following ability |
Parameter Comparison
The visualization compares across versions:
Learning Platform
The tutorial homepage offers four modules: architecture visualization, data-flow animation, evolution timeline, and parameter comparison — highlighting LLaMA's decoder-only text-generation focus, large-scale pretraining, and its role in advancing the open-source AI community.