• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
October 7, 2026
‘Our Team Consists of True Leaders in Their Respective Academic Disciplines
The HSE International Centre of Decision Choice and Analysis studies a wide range of methods for analysing decision-making and possible scenarios for the development of natural, socio-economic, and political phenomena using various mathematical models. The application of advanced mathematical methods to forecasting helps to prevent negative outcomes and avoid erroneous decisions. The HSE News Service spoke to the centre’s director, Prof. Fuad Aleskerov, about its work.
October 6, 2026
International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod Brings Together Scientists from Russia and Serbia
The International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod’ was held at the Nizhny Novgorod House of Scientists from September 23 to 26. The event was organised by HSE University–Nizhny Novgorod and the Nizhny Novgorod House of Scientists, with the participation of Sberbank and the Institute of Physics Belgrade. The symposium was held for the second time: the first conference took place in 2025 and attracted considerable interest from the academic community.
October 5, 2026
‘The Climate Transition Is Not Necessarily a Limitation for Business
Linara Khadimullina works in the field of low-carbon development. In an interview with the Young Scientists of HSE project, she spoke about why nature is not just a beautiful backdrop, her research on the role of sustainable corporate governance in reducing greenhouse gas emissions, and growing plants as a source of inspiration.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

P. 7757–7764.
Anton R., Mikhalchuk M., Rahmatullaev T., Goncharova E., Druzhinina P., Oseledets I., Kuznetsov A.

We introduce methods to quantify how Large Language Models (LLMs) encode and store contextual information, revealing that tokens often seen as minor (e.g., determiners, punctuation) carry surprisingly high context. Notably, removing these tokens — especially stopwords, articles, and commas — consistently degrades performance on MMLU and BABILong-4k, even if removing only irrelevant tokens. Our analysis also shows a strong correlation between contextualization and linearity, where linearity measures how closely the transformation from one layer’s embeddings to the next can be approximated by a single linear mapping. These findings underscore the hidden importance of “filler” tokens in maintaining context. For further exploration, we present LLM-Microscope, an open-source toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions (via an adapted Logit Lens), and measures the intrinsic dimensionality of representations. This toolkit illuminates how seemingly trivial tokens can be critical for long-range understanding.

Language: English
DOI
Text on another site
Keywords: NLPинтерпретируемостьinterpretabilityLLMбольшие языковые моделиОбработка естественного языка (NLP)

In book

Findings of the Association for Computational Linguistics: NAACL 2025
Association for Computational Linguistics, 2025.
Similar publications
Создание биографического датасета русских прозаиков XX века с помощью ChatGPT-4o
Sherstinova T., Урих А. Е., Вестник Новосибирского государственного университета. Серия: Лингвистика и межкультурная коммуникация 2026 Т. 24 № 2 С. 42–54
Цифровые технологии XXI века способствуют формированию новых методологических подходов в гуманитарных науках, включая литературоведение и биографические исследования. Для автоматизации процессов сбора и аннотирования информации все чаще используют большие языковые модели (Large Language Models, LLM), включая ChatGPT. Развитие подобных инструментов стимулирует переход от традиционных методологических подходов, основанных на исключительно ручном аннотировании и экспертной верификации, к гибридным, ...
Added: September 28, 2026
JavaCapsule: Итеративная генерация и отладка Java-кода на основе структурированной обратной связи
Alexandrov D., Vasilevsky V., Rezunik L. et al., В кн.: Труды Института системного программирования РАНТ. 38. Вып. 5.: ИСП РАН, 2026. С. 73–90.
Ограничения при выполнении задач, требующих точного извлечения информации из сверхдлинных последовательностей, приводят к снижению качества генерации и отладки кода крупных объектно-ориентированных систем, характеризующихся сложными межклассовыми зависимостями и необходимостью анализа длинных трассировок выполнения, превышающих размеры стандартных контекстных окон. Предлагается технология JavaCapsule для генерации и отладки Java-кода, реализующая итеративную отладку на основе структурированной обратной связи, включающей результаты ...
Added: September 24, 2026
Critical Success Factors for GenAI Digital Products: From Technology–Market Duality to a Technology–Market–Ethics Triad
Sorokin I., Tekic Z., Technology in Society 2026 Vol. 88 Article 103456
Generative AI (GenAI) digital products challenge established new product development and product success frameworks because performance after launch remains variable: outputs differ across runs and contexts, behavior can shift under drift, and realized value depends on user interaction, oversight, and governance at scale. Focusing on large language model (LLM)-based GenAI digital products, this study identifies ...
Added: September 23, 2026
Convivial or Manipulative Tools? LLMs in Russian Higher Education (The Case of the School of Advanced Studies)
Shipovalova L., Asya A. Filatova, Социология власти 2026 Vol. 38 No. 1 P. 224–243
This article examines the integration of large language model based chatbots into educational practices in Russian higher education through the lens of critical theory and curriculum ideologies. Drawing on Ivan Illich’s concept of convivial tools and Jürgen Habermas’s distinction between instrumental and communicative rationality, the authors conceptualize large language models as technologies capable of operating ...
Added: September 18, 2026
An LLM-Based Approach for Creating Multi-agent Systems
Rezunik L., Alexandrov D., Mikhail Prozorskiy, , in: Intelligent Decision Technologies. Proceedings of the 17th KES-IDT 2025 ConferenceVol. 450.: Cham: Springer, 2026. P. 81–91.
Multi-Agent Systems (MAS) can benefit from Large Language Models (LLMs), but hallucinations pose risks to decision-making. This paper introduces an approach for creating MAS based on LLMs and proposes a generalized architecture for such systems. We ensure that reasoning is conducted through predicate logic to minimize errors, and LLMs are exclusively utilized to translate natural ...
Added: September 14, 2026
Искусственный интеллект и профессия переводчика: критический интегративный обзор изменений рынка труда, качества перевода и профессиональной реконфигурации (2022 - 2026)
Egorova I. S., Journal of Employment and Career 2025 Vol. 4 No. 4 P. 27–43
Background. The rapid diffusion of large language models has intensified claims that professional translation is approaching technological redundancy. The available evidence, however, is dispersed across labour economics, computational linguistics, translation stu-dies, industry research, and translator education, and it does not support a uniform conclu-sion for the profession as a whole.Purpose. This article examines how artificial ...
Added: September 8, 2026
Генерация исходного кода с использованием больших языковых моделей: систематический обзор методологии Вайб-кодинг
Джонов А. Т., Avdoshin S. M., Информационные технологии 2026 Т. 32 № 8 С. 421–427
This systematic review presents an analysis of the "Vibe Coding" methodology — a contemporary approach to the iterative software development process using Large Language Models (LLMs). Code generation tools are transforming software development by enabling programmers to formulate tasks and describe the desired behavior of software in natural language, while LLMs generate source code corresponding ...
Added: August 25, 2026
Вознаграждается ли использование технологий генеративного ИИ на российском рынке труда?
Rozhkova K., Roshchin S., Roshchina Y., Вопросы экономики 2026 № 8 С. 29–54
Стремительное развитие генеративного искусственного интеллекта (генИИ) усилило дискуссии о том, как новые технологии трансформируют рынок труда, особенно позиции и задачи, которые еще недавно мог выполнять только человек. Эксперименты показывают, что генеративный ИИ может обеспечить значительный прирост производительности, особенно менее квалифицированных сотрудников, что может приводить к возникновению индивидуальной зарплатной отдачи на рынке труда. Впервые для России ...
Added: August 11, 2026
Improving Differential Equation Solving in Compact Language Models via Activation Steering and Reinforcement Learning
Surkov A., Ignatenko V., Koltsov S., Computers, Materials and Continua 2026 Vol. 88 No. 3 Article 74
Large language models have recently demonstrated promising capabilities in mathematical reasoning; however, their performance on tasks requiring strict symbolic manipulation, such as solving differential equations, remains limited, especially for compact models. In this work, we investigate whether activation steering combined with reinforcement learning can improve the quality of solutions generated by pretrained language models without ...
Added: July 8, 2026
Proceedings of the 4th Workshop on NLP for Music and Audio (NLP4MusA 2026)
Buzaev F., Mullakhmetov R., Bogachev R. et al., Association for Computational Linguistics, 2026.
Playlist generation based on textual queries using large language models (LLMs) is becoming an important interaction paradigm for music streaming platforms. User queries span a wide spectrum from highly personalized intent to essentially catalog-style requests. Existing systems typically rely on non-personalized retrieval/ranking or apply a fixed level of preference conditioning to every query, which can ...
Added: June 22, 2026
B3Emo: Quantifying Affect as a Double-Edged Sword in Strategic LLM Interactions
Stepin A., Mozikov M., Kabanov A. et al., IEEE Access 2026 Vol. 14 P. 48127–48144
The deployment of large language models (LLMs) in interactive roles such as automated negotiators, customer service agents, and strategic partners requires them to handle not only logical tasks but also the socio-emotional dimensions of interaction. In these situations, success often relies on understanding social cues, building trust, and using persuasion effectively. These skills are closely ...
Added: June 16, 2026
Rank‑Turbulence Delta and interpretable approaches to stylometric Delta measures
Dmitry Pronin, Evgeny Kazartsev, Digital Scholarship in the Humanities 2026 Vol. 41 No. 3 P. 1616–1630
This article repositions Burrows’s Delta as a flexible family of distance measures for exploratory and unsupervised stylometry, where interpretability and stability are as important as predictive accuracy. We introduce two probabilistic extensions, Rank-Turbulence Delta and Jensen–Shannon Delta, by reinterpreting uncentred standardized word-frequency vectors as non-negative representations that can be normalized into probability distributions and compared ...
Added: June 4, 2026
Анализ культурных референций в творчестве А. Вознесенского: цифровое исследование имен персоналий
Tyuryakova-Matveeva D., Цифровые гуманитарные исследования 2026 № 1 С. 4–26
The article explores cultural references in the works of Andrei Voznesensky by analyzing the personalities he mentions. A total of 1,678 works were processed, including poetry, prose, and early unpublished poems. NER methods based on Natasha, spaCy, and LLM Grok tools made it possible to study the frequency of mentions of famous people and their ...
Added: May 31, 2026
Optimizing Computational Infrastructure for Large Language Models in Bioinformatics: A Case Study
Beknazarov N., , in: Parallel Computational Technologies, 19th International Conference, PCT 2025, Moscow, Russia, April 8–10, 2025, Revised Selected Papers. (CCIS, volume 2891)Vol. 2891.: Springer, 2026. P. 3–16.
This paper addresses the challenge of efficiently training Large Language Models (LLMs) on large-scale, sparse omics datasets in high-performance computing (HPC) environments. Using over 1000 BED tracks as a representative data source, we propose a method combining interval-based chunked storage, sparse matrix transformation, and parallel data loading, integrated within a PyTorch Lightning training framework. Our ...
Added: May 19, 2026
От неизвестности к прозрачности: обзор технологий объяснимого ИИ (XAI)
Avdoshin S. M., Pesotskaya E. Y., Информационные технологии 2026 Т. 32 № 4 С. 185–194
With the rapid advancement of artificial intelligence, and deep learning in particular, models have emerged that are capable of delivering highly accurate predictions. However, the internal logic of such models remains difficult to interpret—an issue of critical importance, especially in domains where the correctness of an algorithm directly affects high-stakes decision-making. One promising avenue for ...
Added: May 8, 2026
Персонализированная обратная связь на основе искусственного интеллекта: модель для магистратуры гуманитарного профиля
Подболотова М. И., Адамский А. И., Kolachev N. et al., Высшее образование в России 2026 Т. 35 № 4 С. 21–35
The purpose of the article is to present and justify a pedagogical model of personal ized feedback based on large language models (LLM) for the educational process in a human ities-oriented master’s program. The relevance of the study is determined by the objectives of digital transformation of higher education in the Russian Federation, outlined in Presidential Decree No. 474 ...
Added: May 4, 2026
Об идеологических предвзятостях генеративного ИИ: Российско-украинский конфликт в репрезентации ChatGPT
Baysha O., Trofimov V., Российская школа связей с общественностью 2026 № 40 С. 171–191
A growing number of scholars are warning about the dangers of the reproduction by generative AI of socio-political and ideological biases absorbed by models from the texts on which they were trained. If a given model was trained on Western media texts, it may generate narratives that reproduce West centric views of world events. This ...
Added: April 21, 2026
Large Language Models as Political Actors: Cultural Bias and Epistemic Power
Seredkina E., Seletkova G., Mikhailovsky A., Technology and Language 2026 Vol. 7 No. 1 P. 63–79
The rapid diffusion of Large Language Models (LLMs) into socially and politically sensitive domains raises critical questions about the nature and origins of political bias in artificial intelligence. While existing research often treats bias as a technical flaw to be minimized, this article advances a broader philosophical and cultural interpretation of LLM bias as an ...
Added: April 1, 2026
Granular computing-based deep learning for text classification
Behzadidoost R., Mahan F., Izadkhah H., Information Sciences 2024 Vol. 652 Article 119746
Granular computing involves a comprehensive process that encompasses theories, methodologies, and techniques to solve complex problems, rather than being just an algorithm. As the volume of generated data continues to grow rapidly, data-driven problems have become increasingly complex. Although deep learning models have outperformed traditional machine learning models in solving complex problems, there is still room for enhancing their performance. ...
Added: March 12, 2026
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit