seminars:stat:181108
Differences
This shows you the differences between two versions of the page.
| seminars:stat:181108 [2018/10/31 22:14] – created qyu | seminars:stat:181108 [2018/10/31 22:15] (current) – qyu | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| + | <WRAP centeralign>## | ||
| + | |||
| + | <WRAP 70% center> | ||
| + | ^ **DATE: | ||
| + | ^ **TIME: | ||
| + | ^ **LOCATION: | ||
| + | ^ **SPEAKER: | ||
| + | ^ **TITLE: | ||
| + | </ | ||
| + | \\ | ||
| + | |||
| + | <WRAP center box 80%> | ||
| + | <WRAP centeralign> | ||
| + | We describe a new method for visualizing topics, the distributions over | ||
| + | terms that are automatically extracted from large text corpora using latent variable | ||
| + | models. Our method finds significant n -grams related to a topic, which are then | ||
| + | used to help understand and interpret the underlying distribution. Compared with | ||
| + | the usual visualization, | ||
| + | multi-word expressions provide a better intuitive impression for what a topic is | ||
| + | “about.” Our approach is based on a language model of arbitrary length expressions, | ||
| + | for which we develop a new methodology based on nested permutation tests to find | ||
| + | significant phrases. We show that this method outperforms the more standard use of | ||
| + | chi-square and likelihood ratio tests. We illustrate the topic presentations on | ||
| + | corpora of scientific abstracts and news articles. | ||
| + | </ | ||
| + | |||
| + | |||
