python三國演義人物出場統計

完整代碼

開源代碼

統計三國演義人物高頻次數

#!/usr/bin/env python
# coding=utf-8
#e10.4CalThreeKingdoms.py
import jieba
excludes = {"來到","人馬","領兵","將軍","卻說","荊州","二人","不可","不能","如此"}
txt = open("threekingdom.txt", "rb").read()
words  = jieba.lcut(txt)
counts = {}
for word in words:if len(word) == 1:continueelif word == "諸葛亮" or word == "孔明曰":rword = "孔明"elif word == "關公" or word == "云長":rword = "關羽"elif word == "玄德" or word == "玄德曰":rword = "劉備"elif word == "孟德" or word == "丞相":rword = "曹操"else:rword = wordcounts[rword] = counts.get(rword,0) + 1
for word in excludes:del(counts[word])
items = list(counts.items())
items.sort(key=lambda x:x[1], reverse=True) 
for i in range(55):word, count = items[i]print ("{0:<10}{1:>5}".format(word, count))

代碼運行：人物頻率統計

threekingdom.txt  kingdom.py
kou@ubuntu:~/python/file_文本處理$ python3 kingdom.py 
Building prefix dict from the default dictionary ...
Dumping model to file cache /tmp/jieba.cache
Loading model cost 2.446 seconds.
Prefix dict has been built succesfully.
曹操         1348
劉備         1144
孔明          865
關羽          557
呂布          322
張飛          300

詞云圖片

#!/usr/bin/env python
# coding=utf-8import jieba
import wordcloudf = open("threekingdom.txt","rb")
t = f.read()
f.close()
ls = jieba.lcut(t)
txt = " ".join(ls)
w = wordcloud.WordCloud(    font_path = "NotoSerifCJK-Bold.ttc",\width = 1000,height = 700,background_color = "white",\)w.generate(txt)
w.to_file("gr.png")

代碼運行

在這里插入圖片描述

本文來自互聯網用戶投稿，該文觀點僅代表作者本人，不代表本站立場。本站僅提供信息存儲空間服務，不擁有所有權，不承擔相關法律責任。
如若轉載，請注明出處：http://www.pswp.cn/news/382489.shtml
繁體地址，請注明出處：http://hk.pswp.cn/news/382489.shtml
英文地址，請注明出處：http://en.pswp.cn/news/382489.shtml

如若內容造成侵權/違法違規/事實不符，請聯系多彩編程網進行投訴反饋email:809451989@qq.com，一經查實，立即刪除！