?
Beyond Utility: Assessing Topic Modeling on Visitor Experience Discourse in Russian-Language Museum Reviews
The automatic extraction of thematic structure from user-generated content presents a significant challenge for computational linguistics, a challenge acutely magnified when the textual domain shifts from commercial goods to cultural institutions. This chapter investigates the problem of conducting topic modeling on a particularly refractory corpus: Russian-language reviews of museums. While reviews of consumer goods and services are predominantly framed within a discourse of utility and aesthetic functionality (encompassing ease of use, visual appeal, and time efficiency) reviews of cultural institutions operate within a fundamentally different paradigm. Here, the analytical focus migrates from the object's utility to the subject's experience. The reviewer articulates a deeply personal, often phenomenological encounter (an aesthetic, emotional, or sensuous impression) invariably grounded in their own biographical history and worldview. This shift from the practical to the experiential renders standard topic modeling assumptions and evaluation metrics problematic.
To address this, we compare the quality of different topic modeling methods that have been used to extract topics from author-curated dataset of 18,737 museum reviews. For comparison we have chosen two popular and well-established algorithms, LDA and BERTopic, and a newly emerging paradigm that utilizes large language models. The extracted topics are evaluated using a tripartite metric system: a combination of clustering and coherence metrics along with topic diversity, intra- and iter-topic similarity measures.
The metrics reveal strong and weak points of the three methods. LDA demonstrates good clustering ability and provides lexically distinct topics but requires expert-level manual interpretation. BERTopic struggles with both clustering and interpretability in our dataset due to sensitivity to evaluative language and morphological variation in Russian. The LLM-based approach, in contrast, produces semantically coherent and interpretable topics, but at the cost of weaker geometric separation in the embedding space. All in all, LDA and LLM-based topic modeling are viewed as complementary methods for such structurally and thematically versatile material as cultural institution reviews: while LDA is helpful for document clustering and obtaining stable lexical structure of the topics, the LLM-based approach is more useful when the aim is to derive semantically rich, human-readable interpretations of complex user experience.
Qualitative analysis of the obtained topics yields substantive conclusions about museum visitor experience in the Russian context. The extracted topics show that Russian visitors expect not only comfort, accessibility, and service quality, but also opportunities for emotional involvement, spiritual grounding, contact with the national past, and the transmission of historical and patriotic feeling to younger generations.
The contribution of this chapter is threefold: first, we establish a domain-specific benchmark for evaluating topic modeling algorithms on reviews of cultural institutions. By providing a replicable methodology and a curated dataset, we offer a reference point against which future studies employing different models or parameters can be compared. Second, we delineate the salient content-related and genre-specific features of this text type, offering concrete heuristics for the selection of optimal modeling parameters. Third, we provide a description of model-specific drawbacks when applied to studied material. These descriptions can guide both research analysts and managers in choosing the best method for their needs as well as topic modeling researchers in developing better domain-specific topic extraction methodology.