A Visual Grammar Analysis of Multimodal Strategies in Bilibili Food Exploration Video Thumbnails

Authors

  • Zequn Zheng School of International Studies, Hangzhou Normal University, Hangzhou 311100, China Author

DOI:

https://doi.org/10.63313/LLCS.9201

Keywords:

Visual Grammar, Multimodality, Short-video Thumbnails, Food Exploration, Bilibili

Abstract

As a visual interface through which users encounter digital content, short-video thumbnails constitute a significant site for the configuration of multimodal resources and the construction of meaning and visual appeal. Drawing on Kress and van Leeuwen’s (2006) theory of Visual Grammar, this study employs quantitative content analysis to systematically code and statistically examine 115 food exploration video thumbnails produced by six food creators, each recognized as a Bilibili “Top 100 Creator”. The analysis is based on the three dimensions of Visual Grammar: Representational Meaning, Interactive Meaning, and Compositional Meaning. Representational Meaning is examined through action processes, reaction processes, and conceptual representations; Interactive Meaning is analyzed through contact, social distance, and perspective; and Compositional Meaning is investigated through thumbnail type, food proportion, background blurring, and textual overlays. The findings indicate that action processes constitute the dominant form of narrative representation, accounting for 61.74% of the sample. Demand and offer function as the two primary contact types, with discernible variation across individual creators. Medium shots predominate, while eye-level angles constitute the basic visual configuration. Food occupies more than 60% of the visual area in 84.35% of the thumbnails, and textual overlays are present in 78.26% of the sample. Composite thumbnails featuring both human figures and food substantially outnumber food-only thumbnails. These findings suggest that food exploration video thumbnails deploy relatively stable multimodal configurations through the combination of people, food, composition and text, thereby visually representing food and associated experiences. This study demonstrates the applicability of Visual Grammar to short-video thumbnails as an emerging form of digital visual text and provides insights into visual communication practices in food-related short-video content.

References

[1] Al-Ali, M. N., & Hamzeh, M. S. M. (2024). Extra cues extra views: A multimodal detection of Arabic clickbait thumbnail verbo-visual cues. Discourse & Communication, 18(1), 3-27.

[2] Bucher, T., & Helmond, A. (2018). The affordances of social media platforms. The SAGE handbook of social media, 1(1), 233-253.

[3] Dedema, M., & Herring, S. (2023). How cover images represent video content: a case study of Bilibili.

[4] Halliday, M. A. K. (1978). Language as social semiotic: The social interpretation of language and meaning. (No Title).

[5] Halliday, M. A. (1985). An Introduction to Functional Grammar. London: Edward Arnold; 1985. Spoken and written language.

[6] Jewitt, C. (Ed.). (2009). The Routledge handbook of multimodal analysis (Vol. 1). London: Routledge.

[7] Kress, G. R., & Van Leeuwen, T. (1996). Reading images: The grammar of visual design.

[8] Kress, G., & Van Leeuwen, T. (2006). Reading Images: The Grammar of Visual Design, 2nd edn Routledge.

[9] Ledin, P., & Machin, D. (2020). Introduction to multimodal analysis. Bloomsbury Publishing.

[10] Piqueras-Fiszman, B., & Spence, C. (2015). Sensory expectations based on product-extrinsic food cues: An interdisciplinary review of the empirical evidence and theoretical accounts. Food quality and preference, 40, 165-179.

[11] Riyandi, S. W. (2022). Visual and verbal means to attract our clicks: Multimodality in YouTube thumbnails. NOTION: Journal of Linguistics, Literature, and Culture, 4(1), 54-62.

[12] Spence, C. (2015). Multisensory flavor perception. Cell, 161(1), 24-35.

[13] Spence, C., Okajima, K., Cheok, A. D., Petit, O., & Michel, C. (2016). Eating with our eyes: From visual hunger to digital satiation. Brain and cognition, 110, 53-63.

[14] Ranieri, M., & van Dijck, J. The Culture of Connectivity. A Critical History of Social Media.

Downloads

Published

2026-08-28

Issue

Section

Articles

How to Cite

A Visual Grammar Analysis of Multimodal Strategies in Bilibili Food Exploration Video Thumbnails. (2026). Literature, Language and Cultural Studies, 6(2), 83–94. https://doi.org/10.63313/LLCS.9201