Ubuntu系统下用Python进行自然语言处理的方法
时间:2026-08-21 | 作者:清风无痕 | 阅读:01. 安装必要的工具与库
在Ubuntu系统中,首先需要安装Python环境(建议使用Python 3.6+)及NLP所需的库。
打开终端,执行以下命令:
- 更新系统包:
sudo apt update && sudo apt upgrade -y - 安装Python3及pip:
sudo apt install python3 python3-pip -y - 安装常用NLP库:
pip3 install nltk spacy textblob gensim transformers - 下载NLTK资源(首次使用需下载):
python3 -c "import nltk; nltk.download('punkt'); nltk.download('stopwords')" - 下载spaCy英文模型:
python3 -m spacy download en_core_web_sm
2. 文本预处理:NLP的基础步骤
文本预处理是将原始文本转换为适合分析的格式。
主要包括分词、去停用词、词干提取/词形还原。
-
分词:将句子拆分为单词或标记。使用NLTK的
word_tokenize或spaCy的tokenizer:import nltkfrom nltk.tokenize import word_tokenizetext = "Natural Language Processing is fascinating."tokens = word_tokenize(text)# 输出:['Natural', 'Language', 'Processing', 'is', 'fascinating', '.']或使用spaCy:
import spacynlp = spacy.load("en_core_web_sm")doc = nlp(text)tokens = [token.text for token in doc]# 输出同上 -
去停用词:过滤掉无实际意义的常见词(如“is”“the”“and”)。使用NLTK的停用词列表:
from nltk.corpus import stopwordsstop_words = set(stopwords.words('english'))filtered_tokens = [word for word in tokens if word.lower() not in stop_words]# 输出:['Natural', 'Language', 'Processing', 'fascinating', '.'] -
词干提取/词形还原:将单词还原为词根形式(如“running”→“run”)。使用NLTK的
PorterStemmer或spaCy的词形还原:from nltk.stem import PorterStemmerstemmer = PorterStemmer()stemmed_tokens = [stemmer.stem(word) for word in filtered_tokens]# 输出:['natur', 'languag', 'process', 'fascin', '.']# spaCy词形还原lemmatized_tokens = [token.lemma_ for token in doc]# 输出:['natural', 'language', 'processing', 'be', 'fascinate', '.']
3. 关键NLP任务实现
(1)词性标注(POS Tagging)
识别文本中每个单词的词性,如名词、动词、形容词。
可以使用NLTK或spaCy:
# NLTKimport nltkfrom nltk.tokenize import word_tokenizefrom nltk import pos_tagtokens = word_tokenize(text)pos_tags = pos_tag(tokens)# 输出:[('Natural', 'JJ'), ('Language', 'NNP'), ('Processing', 'NNP'), ('is', 'VBZ'), ('fascinating', 'JJ'), ('.', '.')]
# spaCyfor token in doc:print(token.text, token.pos_)# 输出:Natural NOUN, Language PROPN, Processing PROPN, is AUX, fascinating ADJ, . PUNCT
(2)命名实体识别(NER)
识别文本中的实体,如人名、地名、组织名。
可以使用NLTK或spaCy:
# NLTKimport nltkfrom nltk import ne_chunk, pos_tag, word_tokenizetokens = word_tokenize(text)pos_tags = pos_tag(tokens)ner_tree = ne_chunk(pos_tags)# 输出:(S (ORGANIZATION Natural) (ORGANIZATION Language) (ORGANIZATION Processing) is fascinating .)
# spaCyfor ent in doc.ents:print(ent.text, ent.label_)# 若文本中有人名/地名,会输出对应实体及类型(如"Apple" → ORG)
(3)情感分析
判断文本的情感倾向,包括积极、消极、中性。
使用TextBlob:
from textblob import TextBlobblob = TextBlob("I love Python. It's amazing!")sentiment = blob.sentiment# 输出:Sentiment(polarity=0.8, subjectivity=0.75)# polarity范围[-1,1],>0为积极,<0为消极;subjectivity范围[0,1],>0.5为主观
(4)主题建模(LDA)
发现文本中的隐藏主题。
使用Gensim:
from gensim import corpora, modelsfrom nltk.tokenize import word_tokenizefrom nltk.corpus import stopwords# 准备语料库texts = ["Python is great for data analysis.", "Data science requires Python skills.", "Machine learning uses Python libraries."]tokens = [word_tokenize(text.lower()) for text in texts]stop_words = set(stopwords.words('english'))filtered_tokens = [[word for word in token if word.isalpha() and word not in stop_words] for token in tokens]# 创建词典和语料库dictionary = corpora.Dictionary(filtered_tokens)corpus = [dictionary.doc2bow(text) for text in filtered_tokens]# 训练LDA模型(2个主题)lda_model = models.LdaModel(corpus=corpus, id2word=dictionary, num_topics=2, passes=10)topics = lda_model.print_topics()# 输出每个主题的关键词及权重for topic in topics:print(topic)
4. 进阶:使用预训练模型
如果想再往前走一步,就可以直接用预训练模型,比如 BERT。
Hugging Face 的transformers库已经封装好了现成的 BERT 模型。
拿来做文本分类、问答这类任务会方便很多:
from transformers import pipeline# 加载预训练的情感分析模型classifier = pipeline("sentiment-analysis")result = classifier("I love Ubuntu and Python!")# 输出:[{'label': 'POSITIVE', 'score': 0.9998}]
注意事项
- Ubuntu系统需确保Python环境配置正确(建议使用
venv创建虚拟环境); - 大规模文本处理时,spaCy的性能优于NLTK;
- 预训练模型(如BERT)需要更多计算资源,可根据需求选择轻量级模型(如DistilBERT)。
免责声明:文中图文均来自网络,如有侵权请联系删除,心愿游戏发布此文仅为传递信息,不代表心愿游戏认同其观点或证实其描述。
相关文章
更多-
- Ubuntu下使用Git管理Swagger项目的完整指南
- 时间:2026-08-31
-
- Ubuntu下GIMP滤镜使用全指南:从基础操作到批量处理
- 时间:2026-08-31
-
- Ubuntu下使用Telnet进行远程管理的完整指南
- 时间:2026-08-31
-
- Ubuntu中如何查看磁盘I/O对CPU的影响:cpustat替代方案
- 时间:2026-08-31
-
- Ubuntu系统使用cpustat查看多核CPU信息的方法
- 时间:2026-08-31
-
- Ubuntu中Rust日志系统配置指南
- 时间:2026-08-27
-
- Ubuntu下Rust集成第三方库完整指南
- 时间:2026-08-27
-
- Ubuntu下Rust跨平台开发:从环境搭建到自动化测试指南
- 时间:2026-08-27
精选合集
更多大家都在玩
大家都在看
更多-
- 以下哪种食材被称为“地下苹果” 蚂蚁庄园今日答案9.18
- 时间:2026-09-17
-
- 蚂蚁庄园今天答题答案2026年9月18日
- 时间:2026-09-17
-
- 蚂蚁庄园答题今日答案2026年9月18日
- 时间:2026-09-17
-
- 蚂蚁庄园小课堂2026年9月18日最新题目答案
- 时间:2026-09-17
-
- 小鸡答题今天的答案是什么2026年9月18日
- 时间:2026-09-17
-
- 蚂蚁庄园每日答题答案2026年9月18日
- 时间:2026-09-17
-
- 蔬菜洗完掉色,说明是被染色了,是真的吗 蚂蚁庄园今日答案9月18日
- 时间:2026-09-17
-
- 满襟蜡绘花纹巧染就花纹当绣裳说的是哪种传统技艺 蚂蚁新村今日答案2026.9.17
- 时间:2026-09-17
