A heuristic approach to determine an appropriate number of topics in topic modeling
关键词: Rate of perplexity change (RPC)perplexitytopic numberlatent Dirichlet allocation (LDA)
英文摘要: Abstract(#br) Background(#br)Topic modelling is an active research field in machine learning. While mainly used to build models from unstructured textual data, it offers an effective means of data mining where samples represent documents, and different biological endpoints or omics data represent words. Latent Dirichlet Allocation (LDA) is the most commonly used topic modelling method across a wide number of technical fields. However, model development can be arduous and tedious, and requires burdensome and systematic sensitivity studies in order to find the best set of model parameters. Often, time-consuming subjective evaluations are needed to compare models. Currently, research has yielded no easy way to choose the proper number of topics in a model beyond a major iterative...
