TY - GEN
T1 - A Text is Worth Several Tokens
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Nie, Zhijie
AU - Zhang, Richong
AU - Wu, Zhanyu
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Text embeddings from large language models (LLMs) have achieved excellent results in tasks such as information retrieval, semantic textual similarity, etc. In this work, we show an interesting finding: when feeding a text into the LLM-based embedder, the obtained text embedding can be aligned with the key tokens in the input text. We first fully analyze this phenomenon on eight LLM-based embedders and show that this phenomenon is universal and is not affected by model architecture, training strategy, and embedding method. Upon further analysis, we find that the main change in embedding space between these embedders and their LLM backbones lies in the first principal component. By adjusting the first principal component, we can align text embedding with the key tokens. Finally, we demonstrate the broad application potential of this finding: (1) we propose a simple and practical sparse retrieval method based on the aligned tokens, which can achieve 80% of the dense retrieval effect of the same model while reducing the computation significantly; (2) we show that our findings provide a novel perspective to help understand novel technologies (e.g., instruction-following embedding) and fuzzy concepts (e.g., semantic relatedness vs. similarity) in this field.
AB - Text embeddings from large language models (LLMs) have achieved excellent results in tasks such as information retrieval, semantic textual similarity, etc. In this work, we show an interesting finding: when feeding a text into the LLM-based embedder, the obtained text embedding can be aligned with the key tokens in the input text. We first fully analyze this phenomenon on eight LLM-based embedders and show that this phenomenon is universal and is not affected by model architecture, training strategy, and embedding method. Upon further analysis, we find that the main change in embedding space between these embedders and their LLM backbones lies in the first principal component. By adjusting the first principal component, we can align text embedding with the key tokens. Finally, we demonstrate the broad application potential of this finding: (1) we propose a simple and practical sparse retrieval method based on the aligned tokens, which can achieve 80% of the dense retrieval effect of the same model while reducing the computation significantly; (2) we show that our findings provide a novel perspective to help understand novel technologies (e.g., instruction-following embedding) and fuzzy concepts (e.g., semantic relatedness vs. similarity) in this field.
UR - https://www.scopus.com/pages/publications/105021055707
U2 - 10.18653/v1/2025.acl-long.379
DO - 10.18653/v1/2025.acl-long.379
M3 - 会议稿件
AN - SCOPUS:105021055707
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 7683
EP - 7694
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -