Auteurs
Gilles de Hollander
Publicatiedatum
2010/6/25
Instituut
Bsc thesis, University of Amsterdam, 2010. http://politicalmashup. nl/2010/12/scriptie-gilles-de-hollander
Beschrijving
In this study parsimonious language models were used to construct word clouds of the proceedings of the European Parliament. Multiple design choices had to be made and are discussed. Important features are stemming during tokenization, including bigrams into the word cloud and multi-lingualism. Also, the original parsimonious language models were extended with an additional term dampening unigrams that already occurred in the word cloud.
This algorithm was tested in a small user study, using proceedings of the Science faculty’s student council. Members of this council had to give their preference for multiple word clouds constructed using either parsimonious language models or simple TF with stop words. 68% over 29%(p< 0.05, two-tailed paired t-test) preferred the word clouds constructed using parsimonious language models.
Totaal aantal citaties