词袋模型、TF-IDF 与文本表示
先计数,后思考。TF-IDF 在定义明确的任务上依然胜过嵌入模型,即便在 2026 年。
类型: 构建
语言: Python
前置条件: Phase 5 · 01 (文本处理),Phase 2 · 02 (从零实现线性回归)
用时: ~75 分钟
问题所在
模型需要数字,而你只有字符串。
每个 NLP 管道都必须回答同一个问题:如何将变长的 token 流转化为分类器可以消费的固定大小向量。这个领域给出的第一个答案是能用的最笨方案——数词频,构造成向量。
这个向量承载了比任何嵌入模型都多的生产级 NLP 应用。垃圾邮件过滤器、主题分类器、日志异常检测、搜索排序(在 BM25 之前)、第一波情感分析、学术界 NLP 基准测试的头十年。2026 年的从业者在窄分类任务上仍然会首先选择它。它快速、可解释,在词的出现与否是关键因素的任务上,往往与 4 亿参数的嵌入模型难以区分。
本课从零构建词袋模型和 TF-IDF,然后展示 scikit-learn 如何用三行代码完成同样的工作,最后指出让你转向嵌入模型的失败模式。
核心概念
词袋模型 (Bag of Words, BoW) 丢弃了词序。对每个文档,统计每个词汇表中词出现的次数。向量长度等于词汇表大小,位置 i 存放词 i 的计数。
TF-IDF 对 BoW 进行重新加权。出现在每个文档中的词信息量低,所以缩小其权重;在整个语料库中稀少但在单个文档中频繁出现的词是信号,所以放大其权重。
TF-IDF(w, d) = TF(w, d) * IDF(w)
= count(w in d) / |d| * log(N / df(w))
其中 TF 是文档中的词频,df 是文档频率(包含该词的文档数量),N 是文档总数。log 使无处不在的词的权重保持有界。
关键特性:两者都产生具有可解释轴的稀疏向量。你可以查看训练好的分类器的权重,读出哪些词将文档推向各个类别。你无法用 768 维的 BERT 嵌入做到这一点。
动手构建
步骤 1:构建词汇表
def build_vocab(docs):
vocab = {}
for doc in docs:
for token in doc:
if token not in vocab:
vocab[token] = len(vocab)
return vocab
输入:已分词的文档列表(任何词级别分词器都可以;本课的 code/main.py 使用简化的纯小写版本)。输出:{word: index} 字典。稳定的插入顺序意味着索引 0 对应第一个文档中第一个出现的词。惯例各不相同;scikit-learn 按字母排序。
步骤 2:词袋模型
def bag_of_words(docs, vocab):
matrix = [[0] * len(vocab) for _ in docs]
for i, doc in enumerate(docs):
for token in doc:
if token in vocab:
matrix[i][vocab[token]] += 1
return matrix
>>> docs = [["cat", "sat", "on", "mat"], ["cat", "cat", "ran"]]
>>> vocab = build_vocab(docs)
>>> bag_of_words(docs, vocab)
[[1, 1, 1, 1, 0], [2, 0, 0, 0, 1]]
行是文档,列是词汇索引。条目 [i][j] 表示"词 j 在文档 i 中出现了多少次"。文档 1 中 cat 出现了两次,所以是 2。文档 0 中 ran 没出现,所以是 0。
步骤 3:词频与文档频率
import math
def term_frequency(doc_bow, doc_length):
return [c / doc_length if doc_length else 0 for c in doc_bow]
def document_frequency(bow_matrix):
df = [0] * len(bow_matrix[0])
for row in bow_matrix:
for j, count in enumerate(row):
if count > 0:
df[j] += 1
return df
def inverse_document_frequency(df, n_docs):
return [math.log((n_docs + 1) / (d + 1)) + 1 for d in df]
两个值得指出的平滑技巧。(n+1)/(d+1) 避免 log(x/0)。尾部的 +1 确保出现在每个文档中的词 IDF 仍为 1(而非 0),与 scikit-learn 的默认行为一致。其他实现使用原始的 log(N/df)。两种方式都可以;平滑版本更友好。
步骤 4:TF-IDF
def tfidf(bow_matrix):
n_docs = len(bow_matrix)
df = document_frequency(bow_matrix)
idf = inverse_document_frequency(df, n_docs)
out = []
for row in bow_matrix:
length = sum(row)
tf = term_frequency(row, length)
out.append([tf_j * idf_j for tf_j, idf_j in zip(tf, idf)])
return out
>>> docs = [
... ["the", "cat", "sat"],
... ["the", "dog", "sat"],
... ["the", "cat", "ran"],
... ]
>>> vocab = build_vocab(docs)
>>> bow = bag_of_words(docs, vocab)
>>> tfidf(bow)
三个文档,五个词汇词(the、cat、sat、dog、ran)。the 出现在全部三个文档中,所以 IDF 低。dog 只出现在一个文档中,所以 IDF 高。向量是稀疏的(大多数条目很小),区分性强的词凸显出来。
步骤 5:L2 归一化
def l2_normalize(matrix):
out = []
for row in matrix:
norm = math.sqrt(sum(x * x for x in row))
out.append([x / norm if norm else 0 for x in row])
return out
没有归一化,较长的文档会得到更大的向量,在相似度计算中占主导。L2 归一化将每个文档映射到单位超球面上。行之间的余弦相似度现在就是点积。
实际应用
scikit-learn 提供了生产级版本。
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
docs = ["the cat sat on the mat", "the dog sat on the mat", "the cat ran"]
bow_vectorizer = CountVectorizer()
bow = bow_vectorizer.fit_transform(docs)
print(bow_vectorizer.get_feature_names_out())
print(bow.toarray())
tfidf_vectorizer = TfidfVectorizer()
tfidf = tfidf_vectorizer.fit_transform(docs)
print(tfidf.toarray().round(3))
CountVectorizer 一次调用完成分词、词汇构建和 BoW。TfidfVectorizer 在此基础上加入 IDF 加权和 L2 归一化。两者都返回稀疏矩阵。对于 10 万文档,密集版本无法放入内存;保持稀疏直到分类器要求密集格式。
改变一切的关键参数:
| ngram_range=(1, 2) | 包含 bigram。通常提升分类效果。 |
| min_df=2 | 丢弃出现在少于 2 个文档中的词。在噪声数据上缩减词汇表。 |
| max_df=0.95 | 丢弃出现在超过 95% 文档中的词。无需硬编码列表即可近似移除停用词。 |
| stop_words="english" | scikit-learn 内置停用词表。取决于任务——情感分析不应丢弃否定词。 |
| sublinear_tf=True | 使用 1 + log(tf) 代替原始 tf。当一个词在某个文档中重复多次时有帮助。 |
TF-IDF 仍然胜出的场景(截至 2026 年)
- 垃圾邮件检测、主题标注、日志异常标记。词的出现是关键;语义细微差别不重要。
- 低数据场景(数百个标注样本)。TF-IDF 加逻辑回归没有预训练成本。
- 任何延迟敏感的场景。TF-IDF 加线性模型在微秒级响应。通过 Transformer 嵌入文档需要 10-100ms。
- 需要解释预测结果的系统。检查分类器系数,排名靠前的正向词就是原因。
TF-IDF 失败的场景
语义盲区失败。看这两个文档:
- “The movie was not good at all.”
- “The movie was excellent.”
一条是负面评价,一条是正面评价。它们的 TF-IDF 重叠部分只有 {the, movie, was}。词袋分类器必须记住 not 出现在 good 附近会翻转标签。在足够多的数据上它可以学到,但永远不如理解句法的模型来得优雅。
另一个失败:推理时的词汇外 (out-of-vocabulary) 词。在 IMDb 评论上训练的 BoW 模型对 Zoomer-approved 这种训练中从未出现的 token 完全无能为力。子词嵌入(第 04 课)可以处理这个问题。TF-IDF 不行。
混合方案:TF-IDF 加权嵌入
2026 年中等数据量分类的实用默认方案:将 TF-IDF 权重作为词嵌入上的注意力权重。
def tfidf_weighted_embedding(doc, tfidf_scores, embedding_table, dim):
vec = [0.0] * dim
total_weight = 0.0
for token in doc:
if token not in embedding_table or token not in tfidf_scores:
continue
weight = tfidf_scores[token]
emb = embedding_table[token]
for i in range(dim):
vec[i] += weight * emb[i]
total_weight += weight
if total_weight == 0:
return vec
return [v / total_weight for v in vec]
你从嵌入获得语义能力,从 TF-IDF 获得稀有词的强调。分类器在池化后的向量上训练。在大约 5 万个标注样本以下的情感分析、主题分类和意图分类任务上,这种方案优于单独使用任一方法。
交付
保存为 outputs/prompt-vectorization-picker.md:
—
name: vectorization-picker
description: Given a text-classification task, recommend BoW, TF-IDF, embeddings, or a hybrid.
phase: 5
lesson: 02
—
You recommend a text-vectorization strategy. Given a task description, output:
1. Representation (BoW, TF-IDF, transformer embeddings, or a hybrid). Explain why in one sentence.
2. Specific vectorizer configuration. Name the library. Quote the arguments (`ngram_range`, `min_df`, `max_df`, `sublinear_tf`, `stop_words`).
3. One failure mode to test before shipping.
Refuse to recommend embeddings when the user has under 500 labeled examples unless they show evidence of semantic failure in a TF-IDF baseline. Refuse to remove stopwords for sentiment analysis (negations carry signal). Flag class imbalance as needing more than a vectorizer change.
Example input: "Classifying 30k customer support tickets into 12 categories. Most tickets are 2-3 sentences. English only. Need explainability for audit logs."
Example output:
– Representation: TF-IDF. 30k examples is not small; explainability requirement rules out dense embeddings.
– Config: `TfidfVectorizer(ngram_range=(1, 2), min_df=3, max_df=0.95, sublinear_tf=True, stop_words=None)`. Keep stopwords because category keywords sometimes are stopwords ("not working" vs "working").
– Failure to test: verify `min_df=3` does not drop rare category keywords. Run `get_feature_names_out` filtered by class and eyeball.
练习
关键术语
| BoW | 词频向量 | 一个文档中词汇表词的计数。丢弃了词序。 |
| TF | 词频 | 一个词在文档中的出现次数,可选择按文档长度归一化。 |
| DF | 文档频率 | 至少包含该词一次的文档数量。 |
| IDF | 逆文档频率 | 经平滑的 log(N / df)。降低出现在所有地方的词的权重。 |
| 稀疏向量 | 大部分为零 | 词汇表通常有 1 万到 10 万个词;大多数在给定文档中不出现。 |
| 余弦相似度 | 向量夹角 | L2 归一化向量的点积。1 表示相同,0 表示正交。 |
网硕互联帮助中心


评论前必须登录!
注册