Key Moments

Stanford CS547 HCI Seminar | Spring 2026 | Show It or Tell It? Text, Visualization, and Combination

Stanford OnlineStanford Online
Education5 min read57 min video
Jul 31, 2026|403 views|15
Save to Pod

Want to know something specific about what's covered?

We've already dissected every moment. Ask and we will deliver (with timestamps).

TL;DR

More text on visualizations is preferred, even over minimalism, but the impact of text on predictions is minimal, while its potential for signaling bias is significant.

Key Insights

1

Titles on charts receive the longest fixations and are most likely to be recalled, indicating the crucial role of text in information visualization.

2

For blind and low-vision individuals, higher-level text descriptions of charts are opposed, while lower-level descriptions are favored, contrasting with sighted readers' preferences.

3

Contrary to minimalist design principles, studies show participants significantly prefer visualizations with more accompanying text, challenging ideas like Tufte's minimalism.

4

Viewer predictions based on ambiguous visualizations are largely unaffected by text annotations, suggesting visual perception can dominate textual influence in such contexts.

5

Text annotations on visualizations provide a strong signal of potential author bias, even if they don't influence predictions.

6

Modern multimodal language models (MLMs) process text and visuals through architectures like unified embedding or cross-modality attention, mirroring some aspects of human cognitive processing but with unknown implications for human comprehension.

Text is a critical, yet historically underestimated, component of visualization

Information visualization aims to provide insight into data through systematic spatial representations. However, the role of language alongside visualizations has been largely overlooked by the visualization community, as evidenced by early design guidelines that omitted text elements. Research, such as the 'Beyond Memorability' study, has shown that titles and labels receive significant attention during chart encoding and are crucial for recall, directly contradicting the notion that text is secondary to visuals. This underscores that language is not merely supplementary but a fundamental part of how users understand and remember information presented in visualizations.

Balancing text and visuals: attention, salience, and takeaway

Deciding what information to convey through text versus visualization is complex. Studies indicate that when text captions align with the most visually salient parts of a chart, users focus on those prominent features. However, when captions highlight less salient areas, the text can override visual salience, influencing user takeaways more strongly than the visual elements themselves. Conversely, if the text is completely unaligned with visual features, the visual elements regain dominance. This dynamic interplay suggests a sophisticated interaction where text and visuals compete and cooperate to shape user understanding, and their relative influence depends on alignment and salience.

Accessibility needs reveal differing preferences for text augmentation

Research on the accessibility of visualizations for blind and low-vision (BLV) individuals has highlighted divergent preferences for text augmentation compared to sighted readers. BLV participants often oppose high-level text descriptions that provide external context, favoring instead lower-level descriptions that detail specific data points. Sighted readers, conversely, tend to prefer high-level interpretations. This finding is significant because it suggests that what might seem helpful for sighted users could be detrimental or irrelevant for users with different sensory needs, emphasizing the importance of considering diverse user groups in design. The conclusion drawn from this work is that natural language should be considered coequal with visualizations in design.

More text is generally preferred, challenging minimalist design

Experiments involving varying amounts of text on charts have revealed a surprising preference for more text, rather than less. Contrary to minimalist design principles, such as those advocated by Edward Tufte, participants consistently favored visualizations that included titles, annotations, and subtitles over those with minimal text. While a small minority expressed a preference for no text or very little text, the majority indicated that 'more text is better.' This challenges the prevailing wisdom of minimizing visual elements and suggests that ample, relevant textual information can significantly enhance user engagement and comprehension.

Text annotations don't sway predictions but signal author bias

Interestingly, when presented with ambiguous visualizations designed for prediction tasks, text annotations had a minimal impact on viewers' predictions. For example, adding text that described a trend did not significantly change whether participants believed a blue or green data line would be higher in the future. However, these same text annotations strongly signaled potential author bias. This suggests that while text may not alter objective interpretations of data in certain contexts, it plays a crucial role in shaping perceptions of credibility and intent.

Cognitive models offer competing explanations for text-visual interaction

Understanding the cognitive underpinnings of text-visualization interaction involves examining theories like Dual Coding Theory and Cognitive Load Theory. Dual Coding Theory posits that separate visual and verbal systems process information independently and simultaneously, enhancing memory and understanding. In contrast, Cognitive Load Theory suggests that limited working memory capacity can be overloaded by presenting information in multiple modalities, negatively impacting comprehension. The effectiveness of combining text and visuals may depend on task complexity, individual abilities, and design choices, with research still being definitive.

System one vs. system two thinking and visualization interpretation

Applying dual process theory, particularly Kahneman and Tversky's System 1 (automatic, fast) and System 2 (deliberate, effortful) thinking, provides a framework for understanding visualization interpretation. System 1 processing relies on quick, visually-based decisions, which might explain why text annotations don't influence predictions in ambiguous visual scenarios. System 2, requiring more cognitive effort and top-down attention, is necessary for accurate, calculated interpretations. This framework suggests that visualizations might predominantly trigger System 1, making them less susceptible to textual influence unless System 2 is explicitly engaged, a process that may require specific design or user intent.

Multimodal language models process text and visuals through complex architectures

Modern multimodal language models (MLMs) integrate text and visual information using advanced architectures like transformers. These models employ strategies such as unified embedding, where visual and text embeddings are concatenated, or cross-modality attention, which allows information flow between different modalities. Architectures like Flamingo use masked cross-attention to selectively link text to relevant image segments. Research into the internal workings of these models, like the Lava models, suggests that early layers process general visual information, middle layers handle semantic processing relevant to the task, and top layers consolidate the answer. While these MLMs show remarkable capabilities in tasks involving both text and visuals, their complex processing mechanisms and their relationship to human cognitive processing remain areas of active research and speculation.

Common Questions

The main challenge is understanding how to optimally balance and integrate text and visualizations for a given message, as the optimal approach is complex and context-dependent. The interaction between language and visual elements is not yet fully understood.

Topics

Mentioned in this video

More from Stanford Online

View all 106 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free