Document Type
Conference Paper
Publication Date
2026
Publication Title
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Pages
649-667
Conference Name
27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, August 2-5, 2026, Atlanta, Georgia, USA
Abstract
Image memes are a pervasive form of online communication, widely used to convey humor, opinions, and cultural references. Prior work has explored making memes accessible to blind users, primarily through auto-generated descriptive captions. While these approaches improve comprehensibility and sometimes incorporate prosodic or emotional cues, they often fail to capture the humor, narrative structure, and contextual nuances that make memes engaging. We present MemeBuddy, a system that models memes as dialog, generating structured, multi-turn audio representations using role-based speakers. MemeBuddy reinterprets a meme as a conversation between two speakers, integrating extracted meme text with contextual knowledge implicitly inferred by a multimodal LLM (e.g., recognition of common meme templates and cultural references) to convey intent, timing, and implicit meaning through conversational interaction. We evaluate MemeBuddy in a user study with 14 blind participants. Results show that dialog-style meme representations consistently improve engagement and user satisfaction compared to caption-style descriptions, while maintaining comparable comprehension.
Rights
© 2026 ACL
Published under the terms of a Creative Commons Attribution 4.0 International (CC BY 4.0) License.
Original Publication Citation
Bhansali, C., Ashok, V. G., & Lee, H.-N. (2026). MemeBuddy: Dialog-style audio representations for engaging non-visual meme experiences. In J. D. Choi, Y.-N. Chen, K. Funakoshi, & A. Emami (Eds.), Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue (pp. 649-667). Association for Computational Linguistics. https://aclanthology.org/2026.sigdial-1.46/
Repository Citation
Bhansali, C., Ashok, V. G., & Lee, H.-N. (2026). MemeBuddy: Dialog-style audio representations for engaging non-visual meme experiences. In J. D. Choi, Y.-N. Chen, K. Funakoshi, & A. Emami (Eds.), Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue (pp. 649-667). Association for Computational Linguistics. https://aclanthology.org/2026.sigdial-1.46/
ORCID
0000-0002-4772-1265 (Ashok)