• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
July 24, 2026
‘I Like Self-Fulfilling Prophecies
Andrey Vorchik studies happiness, delivers popular science lectures, and believes that science should address social issues as well. In an interview for the Young Scientists of HSE University project, he spoke about how emotions influence decision-making, the Bermuda Triangle formed by the bathroom, refrigerator, and bed, and the ideal formula for education.
July 24, 2026
'Physics Is What the World Is Literally Built On'
Physicist Nina Dzhanayeva, recipient of a Vladimir Potanin Foundation scholarship, focuses her research on nanophotonics. In this interview for the HSE Young Scientists project, she discusses nanowells, scientific intuition, and how physics can help in making frangipane cream puffs.
July 20, 2026
Scientists Create Open Dataset for Studying Concentration
A team of Russian researchers, including scientists from HSE University–St Petersburg, has developed the first open multimodal dataset containing recordings of brain activity, heart function, and video observations to help researchers understand what happens in the human brain during deep concentration. In the future, the dataset could accelerate the development of neural interfaces, rehabilitation technologies, and AI systems. The article has been published in Scientific Data.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements

P. 6287–6310.
Voronov A., Wolf L., Ryabinin M.

Large language models demonstrate a remarkable capability for learning to solve new tasks from a few examples. The prompt template, or the way the input examples are formatted to obtain the prompt, is an important yet often overlooked aspect of in-context learning. In this work, we conduct a comprehensive study of the template format’s influence on the in-context learning performance. We evaluate the impact of the prompt template across 21 models (from 770M to 70B parameters) and 4 standard classification datasets. We show that a poor choice of the template can reduce the performance of the strongest models and inference methods to a random guess level. More importantly, the best templates do not transfer between different setups and even between models of the same family. Our findings show that the currently prevalent approach to evaluation, which ignores template selection, may give misleading results due to different templates in different works. As a first step towards mitigating this issue, we propose Template Ensembles that aggregate model predictions across several templates. This simple test-time augmentation boosts average performance while being robust to the choice of random set of templates.

Language: English
DOI
Text on another site
Keywords: Large Language Models

In book

Findings of the Association for Computational Linguistics: ACL 2024
Association for Computational Linguistics, 2024.
Similar publications
Improving Differential Equation Solving in Compact Language Models via Activation Steering and Reinforcement Learning
Surkov A., Ignatenko V., Koltsov S., Computers, Materials and Continua 2026 Vol. 88 No. 3 Article 74
Large language models have recently demonstrated promising capabilities in mathematical reasoning; however, their performance on tasks requiring strict symbolic manipulation, such as solving differential equations, remains limited, especially for compact models. In this work, we investigate whether activation steering combined with reinforcement learning can improve the quality of solutions generated by pretrained language models without ...
Added: July 8, 2026
COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
Parkina U., Rakhuba M., , in: 39th Conference on Neural Information Processing Systems (NeurIPS 2025).: NeurIPS, 2025. P. 71014–71041.
Recent studies suggest that context-aware low-rank approximation is a useful tool for compression and fine-tuning of modern large-scale neural networks. In this type of approximation, a norm is weighted by a matrix of input activations, significantly improving metrics over the unweighted case. Nevertheless, existing methods for neural networks suffer from numerical instabilities due to their ...
Added: April 29, 2026
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
Zmushko P., Beznosikov A., Takáč M. et al., , in: Volume 267: International Conference on Machine Learning, 13-19 July 2025, Vancouver Convention Center, Vancouver, CanadaVol. 267.: [б.и.], 2025. P. 80708–80739.
With the increase in the number of parameters in large language models, the training process increasingly demands larger volumes of GPU memory. A significant portion of this memory is typically consumed by the optimizer state. To overcome this challenge, recent approaches such as low-rank adaptation (LoRA), low-rank gradient projection (GaLore), and blockwise optimization (BAdam) have ...
Added: November 10, 2025
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Rodionov G., Roman Garipov, Alina Shutova et al., , in: 39th Conference on Neural Information Processing Systems (NeurIPS 2025).: NeurIPS, 2025. P. 46592–46633.
Large Language Models (LLMs) have demonstrated the ability to tackle increasingly complex tasks through advanced reasoning, long-form content generation, and tool use. Solving these tasks often involves long inference-time computations. In human problem solving, a common strategy to expedite work is collaboration: by dividing the problem into sub-tasks, exploring different strategies concurrently, etc. Recent research ...
Added: November 6, 2025
Гендерные различия в игре диктатора: сравнение поведения больших языковых моделей и людей
Parshakov P., Paklina S., Matkin N. et al., Вестник Пермского университета. Серия: Экономика 2026 Т. 21 № 1 С. 42–57
Introduction. Large language Models (LLM) are increasingly being used in social sciences to simulate the behavior of experimental participants and analyze norms of cooperation and justice. However, the question remains whether they are capable of reproducing social asymmetries, including gender differences. Goal. The work aims to test whether LLM reproduces gender differences in the Dictator ...
Added: October 27, 2025
Сравнительный анализ поведения больших языковых моделей и людей в игре «Диктатор»
Shenkman E., Parshakov P., Matkin N., Журнал Новой экономической ассоциации 2026 № 2(71) С. 177–203
This article analyzes the behavior of large language models (LLMs) in the Dictator game; the comparison was carried out against the results of previously conducted laboratory behavioral experiments with human participants. The study examines modifications of the basic game, that include the possibility of taking resources from the opponent as well as the introduction of ...
Added: October 27, 2025
Large Language Model Failures in Higher Education: Causes and Prevention
Andrei A. Ternikov, COMPUTER 2025 Vol. 58 No. 11 P. 74–83
The rapid adoption of artificial intelligence and large language models (LLMs) in higher education presents unique technical challenges. This article examines critical failures in LLM implementation across academic environments and provides practical strategies for successful integration. ...
Added: July 31, 2025
A graph-based approach to closed-domain natural language generation
Фирсанова В. И., Научный результат. Серия: Вопросы теоретической и прикладной лингвистики 2024 Vol. 10 No. 3 P. 135–167
Graph-based Natural Language Processing (NLP) methods have seen significant advancements in recent years with the development of Large Language Models (LLMs) and Retrieval Augmented Generation (RAG). LLMs are sophisticated models that recognize numerous NLP tasks by analyzing the users' natural language instructions called prompts. However, their industrial use is questionable due to such ethical concerns ...
Added: July 12, 2025
Smart Technical Support System Development Using Knowledge Map-Aided Approach
Alexander Suleykin, Peter Panfilov, , in: 2024 6th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA).: NY: IEEE, 2024. P. 455–460.
Added: April 5, 2025
Proceedings of the 28th Conference on Computational Natural Language Learning
Association for Computational Linguistics, 2024.
CoNLL is a conference organized yearly by SIGNLL (ACL’s Special Interest Group on Natural Language Learning), focusing on theoretically, cognitively and scientifically motivated approaches to computational linguistics. This year, CoNLL was held alongside EMNLP 2024. ...
Added: March 11, 2025
Подход к созданию сервиса генерации программного кода мобильных приложений с использованием больших языковых моделей
Резуник Л., Александров Д.В., ИТ-Стандарт 2024 № 4 С. 34–41
Machine learning technologies and various tools for code generation have had a significant impact on the field of software development in recent years. Although most of the existing solutions are not built exactly for code generation, programmers apply them in different tasks. Not many of the existing AI solutions work well with less common languages, ...
Added: December 30, 2024
LLM-KT: A Versatile Framework for Knowledge Transfer from Large Language Models to Collaborative Filtering
Северин Н. Н., Булычев И. Д., Yushkov M. et al., ICDM 2024
We present LLM-KT, a flexible framework designed to enhance collaborative filtering (CF) models by seamlessly integrating LLM (Large Language Model)-generated features. Unlike existing methods that rely on passing LLM-generated features as direct inputs, our framework injects these features into an intermediate layer of any CF model, allowing the model to reconstruct and leverage the embeddings ...
Added: December 13, 2024
ПРИМЕНЕНИЕ СТИЛОМЕТРИИ ДЛЯ ОПРЕДЕЛЕНИЯ СГЕНЕРИРОВАННЫХ ТЕКСТОВ
Е. А. Сальников, А. А. Бонч-Осмоловская, В кн.: Информационные технологии в гуманитарных исследованиях: Материалы Международной научно-практической конференции, Красноярск, 25–28 сентября 2023 г.: Сибирский федеральный университет, 2023. С. 176–182.
В рамках данного доклад будет проанализировано использование стилометрической метрики дельта Бёрроуза в качестве метода для определения искусственного (т. е. сгенерированного языковой моделью) текста. Данными для эксперимента послужили дневники – как дневниковые записи случайно выбранных авторов, так и дневниковые записи М. М. Пришвина. В качестве данных языковых моделей послужили дневниковые записи, сгенерированные при помощи языковых моделей ...
Added: October 11, 2024
Linguacodus: A synergistic framework for transformative code generation in machine learning pipelines
Trofimova E., Emil Sataev, Ustyuzhanin A., PeerJ Computer Science 2024 Vol. 10 Article e2328
In the ever-evolving landscape of machine learning, seamless translation of natural language descriptions into executable code remains a formidable challenge. This paper introduces Linguacodus, an innovative framework designed to tackle this challenge by deploying a dynamic pipeline that iteratively transforms natural language task descriptions into code through high-level data-shaping instructions. The core of Linguacodus is ...
Added: September 27, 2024
Dialogue as Autocommunication - On Interactions with Large Language Models
Kartasheva Anna, Technology and Language 2024 Vol. 5 No. 2 P. 57–66
In a dialog with large language models (LLM) there is a coincidence of the addressee and addressee of the message, so such a dialog can be called autocommunication. A neural network can only answer a question that has a formulation. The question is formulated by the one who asks it, i.e. a human being. Human activity in dialog ...
Added: September 9, 2024
ChatGPT vs. Crowdsourcing vs. Experts: Annotating Open-Domain Conversations with Speech Functions
Ostyakova Lidiia, Smilga V., Petukhova K. et al., , in: Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue.: Prague: Association for Computational Linguistics, 2023. P. 242–254.
This paper deals with the task of annotating open-domain conversations with speech functions. We propose a semi-automated method for annotating dialogs following the topic-oriented, multi-layered taxonomy of speech functions with the use of hierarchical guidelines using Large Language Models. These guidelines comprise simple questions about the topic and speaker change, sentence types, pragmatic aspects of ...
Added: May 24, 2024
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Dettmers T., Ruslan Svirschevski, Vage Egiazarian et al., , in: Proceedings of the 12th International Conference on Learning Representations (ICLR 2024).: ICLR, 2024.
Recent advances in large language model (LLM) pretraining have led to high-quality LLMs with impressive abilities. By compressing such LLMs via quantization to 3-4 bits per parameter, they can fit into memory-limited devices such as laptops and mobile phones, enabling personalized use. However, quantization down to 3-4 bits per parameter usually leads to moderate-to-high accuracy ...
Added: March 5, 2024
Extreme Compression of Large Language Models via Additive Quantization
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev et al., , in: 41st International Conference on Machine Learning, ICML 2024; Vienna; Austria; 21 July 2024 до 27 July 2024.: Maastricht: ML Research Press, 2024.
The emergence of accurate open large language models (LLMs) has led to a race towards quantization techniques for such models enabling execution on end-user devices. In this paper, we revisit the problem of "extreme" LLM compression--defined as targeting extremely low bit counts, such as 2 to 3 bits per parameter, from the point of view ...
Added: March 5, 2024
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit