1. 大语言模型(LLM)核心架构解析
大语言模型(Large Language Model, LLM)的核心架构基于Transformer Decoder,采用自注意力机制和前馈网络等组件构建。这种架构设计源于2017年Google提出的Transformer模型,但经过多年演进已形成独特的技术路线。
1.1 Transformer Decoder架构演进
传统Transformer包含Encoder和Decoder两部分,但现代LLM通常仅采用Decoder部分。这种选择基于三个关键考量:
- 自回归特性:Decoder天然适合文本生成任务
- 计算效率:相比Encoder-Decoder结构,纯Decoder计算量更小
- 训练一致性:预训练和微调阶段保持架构统一
典型LLM的Decoder层包含以下核心组件:
- 自注意力机制(Self-Attention)
- 位置编码(Positional Encoding)
- 前馈网络(Feed-Forward Network)
- 归一化层(Normalization)
- 残差连接(Residual Connection)
1.2 自注意力机制创新
原始Transformer的多头注意力(MHA)在LLM中经历了重要改进:
python复制# 传统MHA实现
class MultiHeadAttention(nn.Module):
def __init__(self, d_model, num_heads):
super().__init__()
self.d_model = d_model
self.num_heads = num_heads
self.d_k = d_model // num_heads
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
`
