T5 (Text-To-Text Transfer Transformer) Model
What is T5?
T5 is a revolutionary pre-trained language model proposed by Google that solves all NLP tasks through a unified text-to-text framework.
Full name: Text-To-Text Transfer Transformer
Core Features
1. Encoder-Decoder Architecture
Combines an encoder and a decoder to achieve powerful sequence-to-sequence learning.2. Unification Philosophy
Frames all NLP tasks as text-to-text conversion problems.3. Pre-training + Fine-tuning
Acquires general language understanding through large-scale pre-training.Why is T5 So Important?
Unified Framework
Unifies classification, translation, summarization, and other tasks into the same format.Strong Performance
Achieved state-of-the-art results on multiple NLP benchmarks.Simplified Development
A unified interface greatly simplifies model deployment and application.The Unification Philosophy
T5's core innovation: unify all NLP tasks into the form "text input → text output"
Task Examples
| Task Type | Input Example | Output Example | |-----------|--------------|----------------| | Text classification | classify sentiment: This product is really great! | positive | | Machine translation | translate English to Chinese: How are you? | 你好吗? | | Summarization | summarize: [long text...] | [summary] | | Question answering | question: What is T5's full name? context: [...] | Text-To-Text Transfer Transformer | | Text generation | generate: Write a poem about spring | [generated poem] |
Task Prefix System
Through unified prefixes, T5 can understand and execute many different NLP tasks:
translate English to French:- English-French translationsummarize:- text summarizationcola sentence:- grammatical acceptabilitystsb sentence1:- semantic similaritymnli premise:- natural language inferencequestion:- question answering- Unified interface - All tasks use the same input/output format, simplifying model design
- Parameter sharing - Different tasks share model parameters, improving training efficiency
- Transfer learning - Pre-trained knowledge transfers easily to downstream tasks
- Multi-task learning - Training multiple tasks simultaneously improves generalization
- Role: Understands the input semantics
- Mechanism: Bidirectional attention for comprehensive understanding of the input
- Advantages:
- Parallel processing for efficient training
- Powerful feature extraction
- Role: Generates the target text
- Mechanism: Autoregressive generation with Cross-Attention
- Advantages:
- Ensures coherent output
- Effectively uses input information
- Suits various generation tasks
- Bidirectional learning - Predicts using both preceding and following context for deeper semantics
- Unsupervised learning - No manual labeling needed; trainable on massive corpora
- Reversible and lossless - Token sequences can be fully restored to the original text
- Strong generalization - Handles out-of-vocabulary (OOV) words well
Advantages of the Unification Philosophy
Encoder-Decoder Architecture
T5 adopts the classic Encoder-Decoder structure.
Encoder
Decoder
Data Flow Process
1. Input tokenization - Input text is tokenized 2. Encoder processing - Self-Attention understands the semantics 3. Encoder-Decoder interaction - The Decoder attends to Encoder outputs via Cross-Attention 4. Decoder generation - Generates target tokens one by one 5. Output - Produces the complete translation/generated result
Pre-training Task: MLM
T5 uses Masked Language Modeling (MLM) for pre-training.
How MLM Works
1. Initialize character vocabulary - Start from basic characters 2. Random masking - Randomly select 15% of tokens to mask 3. Model prediction - Predict masked tokens using surrounding context 4. Iterative merging - Count character-pair frequencies and merge the most common pairs