Correlation analysis of short text based on network model
作者: Dongyang YanKeping LiJingjing Ye
作者单位: 1State Key Laboratory of Rail Traffic Control and Safety, Beijing 100044, China
2School of Electrical Engineering, Beijing Jiaotong University, Beijing 100044, China
刊名: Physica A: Statistical Mechanics and its Applications, 2019, Vol.531
来源数据库: Elsevier Journal
DOI: 10.1016/j.physa.2019.121728
关键词: Long-range correlationNetwork modelShort textFluctuation analysis
原始语种摘要: Abstract(#br)Correlation of words in the text is of great importance in text analysis like text retrieval, keywords extraction, and text clustering. For short text, because of the limited information of text content, it is difficult to catch the correlation well among words. In this paper, we propose an algorithm based on the complex network to calculate the correlation of words in short texts. A new variable Edge-degree is proposed and used in studying the network model of texts. By using fluctuation analysis, we give the condition that Edge-degree correlation between words exists beyond nearest neighbors. Further analysis shows that numerical results of the fluctuation function of Edge-degree act a power law distribution and that the scaling exponent diverges at a long distance under...
