学位论文详细信息
Novel document representations based on labels and sequential information
Representation learning;Topic modeling;Supervised learning;Sequential document modeling;Sentiment analysis;Mood analysis;Matrix factorization;Machine learning;Artificial intelligence
Kim, Seungyeon ; Lebanon, Guy Computer Science Park, Haesun Essa, Irfan Eisenstein, Jacob Bengio, Samy ; Lebanon, Guy
University:Georgia Institute of Technology
Department:Computer Science
关键词: Representation learning;    Topic modeling;    Supervised learning;    Sequential document modeling;    Sentiment analysis;    Mood analysis;    Matrix factorization;    Machine learning;    Artificial intelligence;   
Others  :  https://smartech.gatech.edu/bitstream/1853/53946/1/KIM-DISSERTATION-2015.pdf
美国|英语
来源: SMARTech Repository
PDF
【 摘 要 】

A wide variety of text analysis applications are based on statistical machine learning techniques. The success of those applications is critically affected by how we represent a document. Learning an efficient document representation has two major challenges: sparsity and sequentiality. The sparsity often causes high estimation error, and text's sequential nature, interdependency between words, causes even more complication.This thesis presents novel document representations to overcome the two challenges. First, I employ label characteristics to estimate a compact document representation. Because label attributes implicitly describe the geometry of dense subspace that has substantial impact, I can effectively resolve the sparsity issue while only focusing the compact subspace. Second, while modeling a document as a joint or conditional distribution between words and their sequential information, I can efficiently reflect sequential nature of text in my document representations. Lastly, the thesis is concluded with a document representation that employs both labels and sequential information in a unified formulation.The following four criteria are utilized to evaluate the goodness of representations: how close a representation is to its original data, how strongly a representation can be distinguished from each other, how easy to interpret a representation by a human, and how much computational effort is needed for a representation.While pursuing those good representation criteria, I was able to obtain document representations that are closer to the original data, stronger in discrimination, and easier to be understood than traditional document representations. Efficient computation algorithms make the proposed approaches largely scalable. This thesis examines emotion prediction, temporal emotion analysis, modeling documents with edit histories, locally coherent topic modeling, and text categorization tasks for possible applications.

【 预 览 】
附件列表
Files Size Format View
Novel document representations based on labels and sequential information 3202KB PDF download
  文献评价指标  
  下载次数:10次 浏览次数:10次