顯示具有 theorems and principles 標籤的文章。 顯示所有文章
顯示具有 theorems and principles 標籤的文章。 顯示所有文章

2016年1月12日 星期二

Engelhardt, Y. (2007). Syntactic structures in graphics. Computational Visualistics and Picture Morphology, 5, 23-35.

Engelhardt, Y. (2007). Syntactic structures in graphics. Computational Visualistics and Picture Morphology5, 23-35.

圖形的目的是用來說明某種的資訊,使得原本無法看見的訊息能被看見 (visualizing the nonvisual)。依據過去的文獻,本研究建議將圖形的建構(building blocks)分為三個部分: a) 顯示的圖形物件 (graphic objects)、b) 配置這些物件使其具有意義來說明訊息的圖形空間 (graphic spaces)以及c) 這些物件的圖形性質 (graphic properties)。在此,圖形物件為一個遞迴的結構 (recursive structure),也就是:在一個圖形空間上配置的一組圖形物件能夠共同形成一個更高階的圖形物件。

本研究也提出圖形物件的語法類型 (syntactic categories),用來解釋物件彼此間可以允許的空間關係並且區分圖形的基本構成 (basic constituents)。所有的圖形都是建立在不同語法類型的圖形物件之組合的可能性上,不同語法類型的圖形物件在圖形呈現 (graphic representation)上有不同的行為,此一限制產生它們有不同的空間定位。所有圖形物件的語法類型可分為兩大群組:1) 依附於圖形空間位置上的物件;2) 依附於其他物件上的物件。前者包括節點 (node)、線標示 (line locator)、面標示 (surface locator)以及格標記 (grid marker);後者則有標籤 (label)、連結線 (connector)、比例區段 (proportional segment)和框架 (frame)。以地圖為例,節點 、線標示、面標示可分別代表地圖上城市、河流與湖泊或國家的標示,格標記則用於表示地圖上的經緯度,標籤則是標明節點代表的城市名稱。各種圖形物件的類型、依附類型與例子,如下表:




根據語言的語法學概念,圖形的語法學 (the syntactics of graphics)研究不同語法類型的圖形物件之間的關係,研究圖形物件與圖形空間之間以規則與限制為基礎的關係,也研究圖形物件如何組合成複合的圖形物件以及複合的圖形物件如何能以更簡單的圖形物件分析。圖形可分為實體場景和物件的影像以及抽象的圖形,前者如照片與地圖等呈現實體的空間,後者則如家族樹 (family trees)、統計圖表 (statistical charts) 等呈現概念性的空間。照片與地圖等利用影像中的空間配置呈現出真實的空間配置;家族樹和圓餅圖 (pie charts) 則以影像中的空間配置呈現非空間性的資訊。然而,實體空間的呈現並不一定表達出被呈現物件的真實座標比例(the true coordinate proportions),許多圖形也同時由實體和概念性的空間組合而成,例如在地圖上以高度呈現國家的人口密度分布。下表是若干圖形空間的類型與代表。


Building upon the existing literature, we are suggesting to regard the building blocks of all graphics as falling into three main categories: a) the graphic objects that are shown (e.g., a dot, a pictogram, an arrow), b) the meaningful graphic spaces into which these objects are arranged (e.g., a geographic coordinate system, a timeline), and c) the graphic properties of these objects (e.g., their colors, their sizes).

We suggest that graphic objects come in different syntactic categories, such as nodes, labels, frames, links, etc. Such syntactic categories of graphic objects can explain the permissible spatial relationships between objects in a graphic representation.

In addition, syntactic categories provide a criterion for distinguishing meaningful basic constituents of graphics.

It is about images that can be regarded as ‘visualizing the nonvisual’ in an attempt to clarify information of some sort. Such images are often collectively referred to as “graphics”.

In 1914, Willard Brinton writes in his book Graphic methods for presenting facts that “The principles for a grammar of graphic presentation are so simple that a remarkably small number of rules would be sufficient to give a universal language”.

In 1967, Jacques Bertin publishes his classic Sémiologie graphique, in which he analyses the “language” of graphic representations and the “visual variables of the image”.

In 1976, linguist Ann Harleman Stewart examines the properties of diagrams and claims that “Like any language, graphic representation has a vocabulary and a grammar”.

In 1984, Clive Richards proposes a “grammatically-based analysis” of diagrams in his Ph.D. thesis Diagrammatics.

In 1986, Jock Mackinlay suggests that “graphical presentations are actually sentences of graphical languages that have precise syntactic and semantic definitions”. In Mackinlay’s approach, “the syntax of a graphical language is defined to be a set of well-formed graphical sentences”.

In 1987, Fred Lakin publishes his paper “Visual grammars for visual languages”, in which he describes his approach to the “spatial parsing” of graphics, which he defines as “the process of recovering the underlying syntactic structure of a visual communication object from its spatial arrangement”.

Kress and van Leeuwen publish their book Reading images: the grammar of visual design (1996). Unfortunately, it is difficult to extract a systematic approach to a syntactic analysis of graphics from their book.

A paper titled “The visual grammar of information graphics” (1996) by Engelhardt et al., suggests “syntactic categories of visual components”.

Robert Horn, in his book Visual Language (1998), proposes a morphology and a syntax of visual language based partly on the work of Jacques Bertin and on the Gestalt principles of perception.

In his book The grammar of graphics (1999), Leland Wilkinson describes an approach to graphics that is related to object-oriented design in computer science. However, he uses grammatical terminology “metaphorically”, and not in a linguistic sense.

Colin Ware (2000) writes about the “perceptual syntax of diagrams”, describing “the grammar of node-link diagrams” and “the grammar of maps”.

Engelhardt
, in his Ph.D. thesis The language of graphics (2002) provides a detailed proposal for the analysis of syntactic structure, which he applies to a broad spectrum of graphic representations.

We propose a notion of graphic objects that will allow for recursive structures: Any graphic representation – and any meaningful visible component of a graphic representation – may be referred to as a graphic object. This means that graphic objects can be distinguished at various levels of a graphic representation. For example, a map or a chart in its entirety is a graphic object. In addition, the various symbols or components that are positioned within that map or chart are graphic objects as well.

A bottom-up description of this principle was given above: a set of graphic objects can be arranged into a graphic space, together forming a single graphic object at a higher level. This “nesting” or “embedding” (Engelhardt 2002) of graphic structures can be referred to as “recursive composition” (Card 2003).

In technical terms, a meaningful graphic space could be defined as a graphic space that involves an interpretation function from spatial positions to one or more domains of information values.

In graphics, not only the possible constituents themselves (graphic objects), and the diverse possible ways of arranging these constituents (in meaningful graphic spaces), but also the possible visual appearances of these constituents (graphic properties such as size, color), could be considered as being part of the graphic “vocabulary”. In this sense we can say that the building blocks of graphics fall into three main categories: graphic objects, meaningful graphic spaces, and graphic properties.



To make a more general statement, we claim that all graphics are based on the possibility of combining graphic constituents (graphic objects) of different syntactic categories (Engelhardt et al. 1996, Engelhardt 2002, 2006).

Graphic objects of different syntactic categories “behave” differently in a graphic representation. The constraints that govern their spatial positioning are different.




All syntactic categories of graphic objects can be divided into two main groups: 1) objects that are attached to locations in graphic space (e.g., node, line locator, surface locator, grid marker are all attached to locations in graphic space), and 2) objects that are attached to other objects (label, connector, proportional segment, frame are all attached to other objects).

Richards (1984) believes that “there seems to be little profit in using such items as an individual dot or line as a unit of analysis. If we are going to use linguistics as a model, then what is needed for present purposes is not the pictorial equivalent of a phoneme or morpheme but something closer to a noun phrase”.

The basic graphic objects in a particular graphic representation are those that can be regarded as functioning in some syntactic category within that particular graphic representation (e.g., as a label, as a node, as a connector, as a proportional segment, etc.).

The distinction between syntactics, semantics, and pragmatics was introduced by Charles Morris (1938, 1946). Morris conceives of syntactics as the investigation of the relationships between signs, of the ways in which complex signs can be constructed from simple ones, as well as the ways in which complex signs can be analyzed into more simple ones (Morris 1946/1971).

The syntactics of graphics investigates the relationships between graphic objects of different syntactic categories. It investigates the rule- and constraint-based relationships between graphic objects (of different syntactic categories) and graphic spaces.

And syntactics investigates how graphic objects can be combined into composite graphic objects, and how composite graphic objects can be analyzed into more simple ones.

Looking at the broad spectrum of graphics we can say that images of physical scenes and objects, such as pictures and maps, represent physical spaces, while many abstract graphics, such as family trees and statistical charts, represent conceptual spaces (Engelhardt 1999, 2002).

In other words, pictures and maps use spatial arrangement in the image to represent spatial arrangement in the world, while family trees and pie charts use spatial arrangement in the image to represent non-spatial information.

Representations of physical spaces do, by the way, not always have to express the true co-ordinate proportions of the represented objects.

Many graphics combine physical and conceptual spaces.

As an example of a true hybrid space (Engelhardt 1999, 2002), think of a three-dimensional landscape drawing of a country in which the drawn “mountains” do not represent physical mountains, but – for example - population density, peaking in the cities and flat in the countryside. In this case, the horizontal plane represents the physical space of the country’s geography, while the vertical dimension represents the conceptual space of population density.




We claim that all types of graphic representation of information can be analyzed in terms of their composition from graphic spaces of different sorts.

We have tried to show that specifying such a visual language means a) specifying the syntactic categories of its graphic objects, plus b) specifying the graphic space in which these graphic objects are positioned, plus c) specifying the visual coding rules that determine the graphic properties of these graphic objects (see table 1).

The syntactic structure of a graphic representation is determined by the rules of attachment for each of the involved syntactic categories (see table 2) and by the structure of the meaningful graphic space that is involved (see table 3).

With this analysis we have attempted to demonstrate that Morris’ original notion of syntactics applies well to the structure of graphics.

2016年1月2日 星期六

Borkin, M., Vo, A., Bylinskii, Z., Isola, P., Sunkavalli, S., Oliva, A., & Pfister, H. (2013). What makes a visualization memorable?. Visualization and Computer Graphics, IEEE Transactions on, 19(12), 2306-2315.

Borkin, M., Vo, A., Bylinskii, Z., Isola, P., Sunkavalli, S., Oliva, A., & Pfister, H. (2013). What makes a visualization memorable?. Visualization and Computer Graphics, IEEE Transactions on19(12), 2306-2315.

過去許多視覺化專家認為視覺化的結果不應該包括「圖表上的廢話」(chart junk),而應該越清楚地顯示資料越好,例如Edward Tufte 與Stephen Few[13, 14, 37, 38],心理學的實驗結果也證實了簡單而清楚的視覺化較容易閱讀[11, 24]。但近來也有一些研究者發表「圖表上的廢話」也許能夠增進閱讀者的持久力,讓他們更能努力在了解圖表上,增加對資料的認識與知識[4, 8, 19]。Bateman et al. [4] 的研究指出有裝飾的圖表比素樸的圖表有較佳的記憶,並且在理解上也不會較差,神經生物學則有加上視覺困難 (visual difficulties) 可以增強觀看者理解的假說[8, 19]。除了「圖表上的廢話」之外,圖表的類型、顏色與其他美學因素也會影響人們的觀看、解讀與記憶。本研究利用410張具有圖表的媒體,對261位受試者進行圖表記憶的研究,探討圖表是否與自然景象的影像[20]一樣在不同人之間具有一致性,並且測試什麼樣的視覺化類型與屬性具有較好的可記憶性。

為了進行本研究對各種圖表類型與屬性是否影響圖表的可記憶性,本研究首先參考前人的研究制訂視覺化的分類架構 (taxonomy)。過去的分類架構有根據圖形的知覺模型 (graphical perception models)、視覺與組織配置 (visual and organizational layout) 以及圖形資料的編碼 (graphical data encodings) [6, 12, 35, 33],或是根據視覺化的演算法 [32, 36] ,近年也有以互動型視覺化和其提供的任務作為分類架構 [16, 17, 33, 35]。本研究的分類架構以Harris [15]提出的詞彙 (vocabulary) 為主,強調圖形表示的語法結構與資訊類型 (Englehardt [12]) 以及人類對圖形的知覺 (Cleveland and McGill [11]),將視覺化分為12類,每一類又分為若干次分類,如下表所示:


在分類架構上,另外對這些圖表附加上若干性質(properties)與屬性(attributes),性質有維度數目 (2D或3D)、多重性 (單一、群體、多板塊、組合)、插圖(pictorial)以及時間序列,屬性尚包括是否是黑白、顏色的數目、資料與非資料間的比例 (data-ink ratio)、視覺元素面積佔整個影像的比例、是否有插圖以及是否有人像等。

不同的視覺化影像來源中,科學出版品 (scientific publications) 和資訊圖表 (infographics) 在視覺化上具有的多重性非當高,顯示這兩種影像來源的製作上往往需要組合多種圖表來表示概念和理論,而另一個考量則在於節省篇幅;反之,政府與組織出版品則大多是單一視覺化。

以視覺化的類型而言,為了解釋文章中的概念或是理論的結果,科學出版品有極高的比例使用diagram,此外科學文章中也大量使用基本的視覺編碼技術(visual encoding techniques),諸如折線圖、長條圖和散布圖,某些領域內會使用特定的視覺編碼,例如網格與矩陣圖 (grid and matrix plots)以及樹狀與網路圖 (tree and networks)。因此,相較於其他類型的來源,科學出版品中較'常使用樹狀與網路圖、網格與矩陣圖和散布圖。

資訊圖表中也使用大量的diagrams,主要包括流程圖(flow charts)和時間表(timelines),此外表格(table)也經常被使用在資訊圖表中,這些表格通常會加上插畫的裝飾。但是資訊圖表相較於其他來源,較少使用折線圖。

新聞媒體和政府出版品主要使用長條圖、折線圖地圖和表格,折線圖通常用來表示時間序列的資料。兩種來源的差異,是政府的報告會大量使用圓形圖,例如圓餅圖。

本研究在評估圖表具有可記憶性的高低時,使用命中率 (hit rate, HR) 與誤警率 (false alarm rate, FAR),並且用敏感指數(sensitivity index) d-prime metric整合兩個數值進行排序, d' = Z(HR) − Z(FAR) ,Z 是高斯分布(Gaussian distribution)的累積分配函數(cumulative distribution function, CDF)的倒數(inverse)。愈高的d'值表示HR較高而FAR較低,此時該視覺化圖形較容易被記憶,反之則否。本研究得到視覺化圖形的可記憶性平均測量結果為HR = 55.36%和SD = 16.51%以及FAR = 13.17%, SD = 10.73%。相較於先前的圖像可記憶性的研究結果,較自然景象的可記憶性 (HR = 67.5%, SD = 13.6%以及FAR = 10.7%, SD = 7.6%) [21] 差,但與人臉的可記憶性 (HR = 53.6%, SD = 14.3%以及FAR = 14.5%, SD = 9.9%) [2]差不多。而將本研究的測試結果隨機分為兩群,經過25次的測量,HR、FAR和d-prime在兩群間的結量結果的Spearman等級相關係數分別為0.83、0.78和0.81。由此可知,視覺化圖形在可記憶性的不同受試者測量上也有一致性,可記憶性是視覺化圖形的一種特性。

以下針對視覺化圖形的種類、性質與屬性進行可記憶性的探討;

具有插畫的視覺化圖形明顯比沒有插畫的視覺化圖形有較高的可記憶性,

愈多顏色的視覺化圖形的可記憶性愈佳,甚至在沒有插畫的圖形上,較多顏色的圖形也比只有一種顏色的圖形可記憶性來得高。

視覺元素面積佔整個影像的比例較高的視覺化圖形也比較容易被記憶。

資料與非資料間的比例較低的視覺化圖形的可記憶性較佳。

diagram、網格與矩陣圖以及樹狀與網路圖等類型較具有獨特(unique)呈現的圖形比長條圖、折線圖、表格等較一般(common)呈現的圖形較具有可記憶性,特別是圖形上沒有插圖時,因為這類圖形的呈現相近,彼此干擾,使得命中率較低而誤警率較高。

在各種來源的視覺化圖形裡,資訊圖表類具有最高的可記憶性,其次是科學出版品,最不具可記憶性的圖形來源則是政府與世界組織類。

另外,具有圓形或圓的邊的圖形的可記憶性也較高,作者認為插畫與圓形都是較自然的視覺化作品,因此較容易被記憶。


We ran the largest scale visualization study to date using 2,070 single-panel visualizations, categorized with visualization type (e.g., bar chart, line graph, etc.), collected from news media sites, government reports, scientific journals, and infographic sources. Each visualization was annotated with additional attributes, including ratings for data-ink ratios and visual densities.

Altogether our findings suggest that quantifying memorability is a general metric of the utility of information, an essential step towards determining how to design effective visualizations.

The conventional view, promoted by visualization experts such as Edward Tufte and Stephen Few, holds that visualizations should not include chart junk and should show the data as clearly as possible without any distractors [13, 14, 37, 38]. This view has also been supported by psychology lab studies, which show that simple and clear visualizations are easier to understand [11, 24].

At the other end of the spectrum, researchers have published that chart junk can possibly improve retention and force a viewer to expend more cognitive effort to understand the graph, thus increasing their knowledge and understanding of the data [4, 8, 19].

What researchers agree on is that chart junk is not the only factor that influences how a person sees, interprets, and remembers a visualization. Other aspects of the visualization, such as graph type, color, or aesthetics, also influence a visualization’s cognitive workload and retention [8, 19, 39].

Recent work has shown that memorability of images of natural scenes is consistent across people, suggesting that some images are intrinsically more memorable than others, independent of an individual’s contexts and biases [20].

This is because given limited cognitive resources and time to process novel information, capitalizing on memorable displays is an effective strategy.

Recent large-scale visual memory work has shown that existing categorical knowledge supports memorability for item-specific details [9, 22, 23]. In other words, many additional visual details of the image come for free when retrieving memorable items.

For our research, we first built a new broad taxonomy of static visualizations that covers the large variety of visualizations used across social and scientific domains. These visualization types range from area charts, bar charts, line graphs, and maps to diagrams, point plots, and tables.

We then used these 2,070 visualizations in an online memorability study we launched via Amazon’s Mechanical Turk with 261 participants. This study allowed us to gather memorability scores for hundreds of these visualizations, and determine which visualization types and attributes were more memorable.

Perception Theory and the Chart Junk Debate:

Bateman et al. conducted a study to test the comprehension and recall of graphs using an embellished version and a plain version of each graph [4]. They showed that the embellished graphs outperformed the plain graphs with respect to recall, and the embellished versions were no less effective for comprehension than the plain versions.

There has been some support for the comprehension results from a neurobiological standpoint, as it has been hypothesized that adding “visual difficulties” may enhance comprehension by a viewer [8, 19].

Other studies have shown that the effects of stylistic choices and visual metaphors may not have such a significant effect on perception and comprehension [7, 39].

In response to the Bateman study, Stephen Few wrote a comprehensive critique of their methodology [14], most of which also applies to other studies. A number of these studies were conducted with a limited number of participants and target visualizations. Moreover, in some studies the visualization targets were designed by the experimenters, introducing inherent biases and over-simplifications [4, 8, 39].

Visualization Taxonomies:

Within the academic visualization community there have been many approaches to creating visualization taxonomies.

Traditionally many visualization taxonomies have been based on graphical perception models, the visual and organizational layout, as well as the graphical data encodings [6, 12, 35, 33].

Another approach to visualization taxonomies is based on the underlying algorithms of the visualization and not the data itself [32, 36].

There is also recent work on taxonomies for interactive visualizations and the additional tasks they enable [16, 17, 33, 35].

Outside of the academic community there is a thriving interest in visualization collections for the general public. For example, the Periodic Table for Management [26] present a classification of visualizations with a multitude of illustrated diagrams for business. The online community Visualizing.org introduces an eight-category taxonomy to organize the projects hosted on their site [25]. InfoDesignPatterns.com classifies visualization design patterns based upon visual representation and user interaction [5].

Cognitive Psychology:

These studies have demonstrated that the differences in the memorability of different images are consistent across observers, which implies that memorability is an intrinsic property of an image [21, 20].

Brady et al. [9] tested the long-term memory capacity for storing details by detecting repeat object images when shown pairs of objects, one old and one new. They found that participants were accurate in detecting repeats with minimal false alarms, indicating that human visual memory has a higher storage capacity for minute details than was previously thought.

More recently, Isola et al. have annotated natural images with attributes, measured memorability, and performed feature selection, showing that certain features are good indicators of memorability [20, 21]. Memorability was measured by launching a “Memory Game” on Amazon Mechanical Turk, in which participants were presented with a sequence of images and instructed to press a key when they saw a repeat image in the sequence. The results showed that there was consistency across the different participants, and that people and human-scale objects in the images contribute positively to the memorability of scenes. That work also showed that unusual layouts and aesthetic beauty were not overall associated with high memorability across a dataset of everyday photos [20].


The taxonomy classifies static visualizations according to the underlying data structures, the visual encoding of the data, and the perceptual tasks enabled by these encodings.

It contains twelve main visualization categories and several popular sub-types for each category. In addition, we supply a set of properties that aid in the characterization of the visualizations.

This taxonomy draws from the comprehensive vocabulary of information graphics presented in Harris [15], the emphasis on syntactic structure and information type in graphic representation by Englehardt [12], and the results of Cleveland and McGill in understanding human graphical perception [11].



Dimension represents the number of dimensions (i.e., 2D or 3D) of the visual encoding.

Multiplicity defines whether the visualization is stand-alone (single) or somehow grouped with other visualizations (multiple). We distinguish several cases of multiple visualizations. Grouped means multiple overlapping/superimposed visualizations, such as grouped bar charts; multi-panel indicates a graphic that contains multiple related visualizations as part of a single narrative; and combination indicates a graph with two or more superimposed visualization categories (e.g., a line plot over a bar graph).

The pictorial property indicates that the encoding is a pictogram (e.g., a pictorial bar chart). Pictorial unit means that the individual pictograms represent units of data, such as the Istotype (International System of Typographic Picture Education), a form of infographics based on pictograms developed by Otto Neurath at the turn of the 19th century [29].

Finally, time is included, specifically as a time series, as it is such a common feature of visualizations and dictates specific visual encoding aspects regarding data encoding and ordering.

The first two attributes, black & white and number of distinct colors give a general sense of the amount of color in a visualization.

A measure of chart junk and minimalism is encapsulated in Edward Tufte’s data-ink ratio metric [37], which approximates the ratio of data to non-data elements.

The visual density rates the overall density of visual elements in the image without distinguishing between data and non-data elements.

Finally, we have two binary attributes to identify pictograms, photos, or logos: human recognizable objects and human depiction. We explicitly chose to have a separate category for human depictions due to prior research indicating that the presence of a human in a photo has a strong effect on memorability [21].



There is also a very high percentage of multiple visualizations in the scientific publication category. There are two primary explanations for this observation. First, like infographics, multiple individual visualizations are combined in a single figure in order to visually explain scientific concepts or theories to the journal readers. Second, combining visualizations into a single figure (even if possibly not directly related) saves page count and money.

In contrast, a very high ratio of single visualizations is seen in government / world organizations. These visualizations are usually published one-at-a-time within government reports, and there are no page limits or space issues as with scientific journals.



Scientific publications, for example, have a large percentage of diagrams. These diagrams are primarily used to explain concepts from the article, or illustrate the results or theories. Also included are renderings (e.g., 3D molecular diagrams). The scientific articles also use many basic visual encoding techniques, such as line graphs, bar charts, and point plots. Domain-specific uses of certain visual encodings are evident, e.g., grid and matrix plots for biological heat maps, trees and networks for phylogenic trees, etc.

Infographics also use a large percentage of diagrams. These diagrams primarily include flow charts and timelines. Also included in infographics is a large percentage of tables. These are commonly simple tables or ranked lists that are elaborately decorated and annotated with illustrations. Unlike the other categories, there is little use of line graphs.

In contrast to the scientific and infographic sources, the news media and government sources publish a more focused range of topics, thus employing similar visualization strategies. Both sources primarily rely on bar charts and other “common” (i.e., learnt in primary school) forms of visual encodings such as line graphs, maps, and tables. The line graphs are most commonly time series, e.g., of financial data. One of the interesting differences between the categories include the greater use of circle plots (e.g., pie charts) in government reports.

Looking at specific visualization categories, tree and network diagrams only appear in scientific and infographic publications. This is probably due to the fact that the other publication venues do not publish data that is best represented as trees or networks. Similarly, grid and matrix plots are primarily used to encode appropriate data in the scientific context. Interestingly, point plots are also primarily used in scientific publications. This may be due to either the fact that the data being visualized are indeed best visualized as point plot representations, or it could be due to domain-specific visualization conventions, e.g., in statistics.

Worth noting is the absence of text visualizations from almost all publication venues. The only examples of text based visualizations were observed in the news media. Their absence may be explained by the fact that their data, i.e., text, is not relevant to the topics published by most sources. Another possible explanation is that text visualizations are not as “main stream” in any of the visualization sources we examined as compared to other visualization types.

Performance Metrics:

Workers saw each target image at most 2 times (less than twice if they prematurely exited the game). We measure an image’s hit rate (HR) as the proportion of times workers responded on the second (repeat) presentation of the image. In signal detection terms: HR = HITS/(HITS+MISSES).

We also measured how many times workers responded on the first presentation of the image. This corresponds to workers thinking they have seen the image before, even though they have not. This false alarm rate (FAR) is calculated: FAR = FA/(FA+CR) , where FA is the number of false alarms and CR is the number of correct rejections (the absense of a response).

For performing a relative sorting of our data instances we used the d-prime metric (otherwise called the sensitivity index). This is a common metric used in signal detection theory, which takes into account both the signal and noise of a data source, calculated as: d' = Z(HR) − Z(FAR) (where Z is the inverse cumulative Gaussian distribution). A higher d' corresponds to a signal being more readily detected. Thus, we can use this as a memorability score for our visualizations. A high score will require the HR to be high and the FAR to be low. This will ensure that visualizations that are easily confused for others (high FAR) will have a lower memorability score.

In other words, we have measured how visualizations would be remembered if they were images. We observed a mean HR of 55.36% (SD = 16.51%) and mean FAR of 13.17% (SD = 10.73%).

For comparison, scene memorability has a mean HR of 67.5% (SD = 13.6%) with mean FAR of 10.7% (SD = 7.6%) [21], and face memorability has a mean HR of 53.6% (SD = 14.3%) with mean FAR of 14.5%(SD = 9.9%) [2]. This possibly supports our first hypothesis that visualizations are less memorable than natural scenes.

This demonstrates that there is memorability consistency with scenes, faces, and also visualizations, thus memorability is a generic principle with possibly similar generic, abstract, features.

We also measured the consistency of our memorability scores [2, 21]. By splitting the participants into two independent groups, we can measure how well the memorability scores of one group on all the target images compare to the scores of another group (Fig. 3).

Averaging over 25 such random half-splits, we obtain Spearman’s rank correlations of 0.83 for HR, 0.78 for FAR, and 0.81 for d-prime, the latter of which is plotted in Fig. 3.

This high correlation demonstrates that the memorability of a visualization is a consistent measure across participants, and indicates real differences in memorability between visualizations.

In other words, despite the noise introduced by worker variability and by showing different image sequences to different workers, we can nevertheless show that memorability is somehow intrinsic to the visualizations.

Of our 410 target visualizations, 145 contained either photographs, cartoons, or other pictograms of human recognizable objects (from here on out referred to broadly as “pictograms”). Visualizations containing pictograms have on average a higher memorability score (Mean (M)=1.93) than visualizations without pictograms (M = 1.14,t(297) = 13.67, p < 0.001). This supports our second hypothesis.

Thus not all chart junk is created equal: annotations and representations containing pictograms are across the board more memorable.

Thus an image, or image of a visualization, containing a human recognizable object will be easily recognizable and probably memorable. Due to this strong main effect of pictograms, we examined our results for both the cases of visualizations with and without pictograms. As shown in the left-most panel of Fig. 1, all but one of the overall top most memorable images (as ranked by their d-prime scores) contain human recognizable pictograms.

As shown in Fig. 4, there is an observable trend of more colorful visualizations having a higher memorability score: visualizations with 7 or more colors have a higher memorability score (M = 1.71) than visualizations with 2-6 colors (M = 1.48,t(285) = 3.97, p < 0.001), and even more than visualizations with 1 color or black-and-white gradient (M = 1.18,t(220) = 6.38, p < 0.001).

When we remove visualizations with pictograms, the difference between visualizations with 7 or more colors (M = 1.34) and those that have only 1 color (M = 1.00) remains statistically significant (t(71) = 3.61, p < 0.001).

Considering all the visualizations together, we observed a statistically significant effect of visual density on memorability scores with a high visual density rating of “3” (M = 1.83), i.e., very dense, being greater than a low visual density rating of “1” (M = 1.28,t(115) = 6.08, p < 0.001) as shown in Fig. 5.

We also observed a statistically significant effect of the data-to-ink ratio attribute on memorability scores with a “bad” (M = 1.81), i.e., low data-to-ink ratio, being higher than a “good” rating (M = 1.23,t(208) = 6.92, p < 0.001) as shown in Fig. 6. Note that using a corrected t-test, we also arrive at the results that the 3 levels of data-ink ratio are pairwise significantly different from each-other .

Summarizing all of these attribute results: higher memorability scores were correlated with visualizations containing pictograms, more color, low data-to-ink ratios, and high visual densities.

As shown in Fig. 7, diagrams were statistically more memorable than points, bars, lines, and tables. These trends remain observable even when visualizations with pictograms are removed from the data. Other than some minor ranking differences and addition of the map category, the main difference is in the ranking of the table visualization type, which without pictograms becomes least memorable.

The middle panel of Fig. 1 displays the most memorable visualizations that do not contain pictograms. Why are these visualizations more memorable than the ones in the right-most panel?

To start with, qualitatively viewing the most memorable visualizations, most are high contrast. These images also all have more color, a trend quantitatively demonstrated in Sec. 7.2 to be correlated with higher memorability. As compared to the more subdued less memorable visualizations, the more memorable visualizations are easier to see and discriminate as images.

Another possible explanation is that “unique” types of visualizations, such as diagrams, are more memorable than “common” types of visualizations, such as bar charts. This trend is also evident in Fig. 7 in which grid/matrix, trees and networks, and diagrams have the highest memorability scores.

Examples of these unique types of visualizations are each individual and unique, whereas bar charts and line graphs are uniform with limited variability in their visual encoding methodology. Previously it has been shown that an item is more likely to interfere with another item if it has similar category or subordinate category information, but unique exemplars of objects can be encoded in memory quite well [22]. This supports our findings that show high FAR and low HR for table and bar visualizations, which both have very similar visuals within their category (i.e., all the bar charts look alike).

Another possible explanation is that visualizations like bar and line graphs are just not natural. If image memorability is correlated with the ability to recognize natural, or natural looking, objects then people may see diagrams, radial plots, or heat maps as looking more “natural”.

One common visual aspect of the most memorable visualizations is the prevalence of circles and round edges. Previous work has demonstrated that people’s emotions are more positive toward rounded corners than sharp corners [3]. This could possibly support both the trend of circular features in the memorable images as well as the concept of natural-looking visualizations being more memorable since “natural” things tend to be round.

As shown in Fig. 8, regardless of whether the visualizations did or did not include pictograms, the visualization source with the highest memorability score was the infographic category (M = 1.99,t(147) = 5.96, p < 0.001 when compared to the next highest category, scientific publications with M = 1.48), and the visualization source with the lowest memorability score was the government and world organizations category (M = 0.86,t(220) = 8.46, p < 0.001 when compared to the next lowest category, news media with M = 1.46).

This may be a contributing factor to the observed trend (see Fig. 8) that visualization sources that have non-uniform aesthetics tend to have higher memorability scores than sources with uniform aesthetics. This observation refutes our last hypothesis that visualizations from scientific publications are less memorable. This may also be due to the fact that visualizations in scientific publications have a high percentage of diagrams (Fig. 2), similar to the infographic category.

The results of our memorability experiment show that visualizations are intrinsically memorable with consistency across people.

They are less memorable than natural scenes, but similar to images of faces, which may hint at generic, abstract, features of human memory.

Not surprisingly, attributes such as color and the inclusion of a human recognizable object enhance memorability. And similar to previous studies we found that visualizations with low data-to-ink ratios and high visual densities (i.e., more chart junk and “clutter”) were more memorable than minimal, “clean” visualizations.

It appears that we are best at remembering “natural” looking visualizations, as they are similar to scenes, objects, and people, and that pictorial and rounded features help memorability.

More surprisingly, we found that unique visualization types (pictoral, grid/matrix, trees and networks, and diagrams) had significantly higher memorability scores than common graphs (circles, area, points, bars, and lines).

2014年4月26日 星期六

Rodrigues, J. F., Traina, A. J., & de Oliveira, M. C. F. (2006, July). Reviewing data visualization: an analytical taxonomical study. In Information Visualization, 2006. IV 2006. Tenth International Conference on (pp. 713-720). IEEE.

Rodrigues, J. F., Traina, A. J., & de Oliveira, M. C. F. (2006, July). Reviewing data visualization: an analytical taxonomical study. In Information Visualization, 2006. IV 2006. Tenth International Conference on (pp. 713-720). IEEE.

information visualization

資料視覺化(data visualization)希望當將圖形資訊呈現給使用者時,使用者便能夠立即理解,提供較快速而的資料分析機制。目前已經有許多針對資料視覺化技術的分類架構(classification schemes)提出。Keim[13]以進行視覺化的資料類型(data type)、視覺化技術(visualization technique)以及應用的互動/扭曲技術(interaction/distortion technique applied)為分類架構(taxonomy)的三個面向,來將視覺化技術與系統加以歸類(categorize)。Schneiderman [27]的分類學則則是一個成對的系統(a pair wise system),包含分析的資料類型以及一組分析的任務,任務例如概觀(overview)、放大縮小(zoom)、過濾(filter)、選取詳細訊息(details-on-demand)等等。Chi [4]的分類法根據一個特定的視覺化模型中不同的性質,詳細地說明視覺化技術,包含資料、抽象(abstraction)、轉換和映射(transformation和mapping)、呈現(presentation)以及互動(interaction)。Tory and Möoller [29]根據視覺塑模(visual modeling)的直覺知覺,將科學視覺化(scientific visualization)和資訊視覺化(information visualization)分別定義為連續及離散類型 。Wiss and Carr [34]在探討3-D技術時則提出了一個考慮注意(attention)、抽象化(abstraction)以及互動支持性(interaction affordance)等的認知為基礎的分類法 。這些分類架構大多僅針對視覺化過程(visualization process)中的某一面向,然而還存在許多有待解決的問題,例如:視覺探索場景(visual exploration scene)的基礎構件(building blocks)是什麼?互動機制如何連結到事實?

本研究根據視覺化技術的機制和視覺知覺理論(visual perception theory),以解釋視覺化場景如何構成以及它們的構成部分如何有助於視覺理解,做為分析模型。本研究認為視覺化是利用空間化(spatialization)的方式讓資料得以空間性的被感知,並以位置(position)、形狀(shape)和顏色(color)等視覺刺激(visual stimuli) 來代表資料項目或特徵,將可以注意到的差異(noticeable differences)極大化(Ware [32]),空間化後的資料才能夠在知覺注意(conscious attention)前確認事物,能夠立即與容易地引發注意,提昇與加速資料理解。Figure 1表現出這樣的分類法概念。



進一步而言,資料的空間化是將資料從難以說明的原始形式轉換為可視的空間形式,空間化的程序有結構展現(structure exposition)、投射(projection)、樣式定位(patterned positioning)以及 再製(reproduction)等。結構展現是指資料可以嵌入階層或關係網絡等內在結構(intrinsic structures),可以大抵體現出資料的意義,也就是使得相關的資料結構能夠視覺化地覺察,例如樹狀地圖(treemaps) [28] 或力導向圖式布局(force directed graph structures) [8] 都屬於結構展現。樣式定位是根據一個或多個方向、直線、圓,甚至特定模式,依序安排個別資料項目的設定,意圖填滿整個投射空間,有時稱之為密集向素呈現(dense pixel displays),這樣的空間化技術有像素長條圖(pixel bar charts) [15],圓餅圖(pie charts)也屬於這類。在投射這類空間化技術的呈現上,資料項目的位置由一個已知或內涵的數學函數來定義,著名的案例有平行座標(parallel coordinates)與星狀座標(star coordinates) [11]。在再製(reproduction)的呈現上,資料的定位是事先已知的,並且從資料蒐集的系統來決定。

在視覺化場景中,位置是前注意知覺(pre-attention perception)的主要部分,而且它和空間化非常有關。位置包括排列(arrangement)、對應(correspondence)、參考(referential)等三種型式,分別由空間化的各種機制所產生。結構展現產生排列,特定的排列能描述結構、階層或其他整體特性(global property)等資訊,資訊可以從局部的個別元素間相互位置或是整體的景觀概況上察覺,例如樹狀地圖表現資料項目間的階層關係。樣式定位會產生位置上的對應,資料項目的位置決定它的對應特性。投射與再製運用函數或明確的參考來源(explicit referewnce)決定資料的位置,提供使用者得以了解。

形狀可以表達的資訊,包含差異(differentiation)、對應(correspondence)、意義(meaning)和關係(relationship)等。差異是指以呈現的形狀區別這些項目,來提供近一步的詮釋。對應則是利用每一種引人注目的形狀對應一種資訊特徵,例如大小比例便是常用的方式。意義是指呈現的形狀本身便帶有意義,這種方式依賴使用者的知識與過去經驗,例如箭頭。關係:線條、外廓或表面可以利用來表現一組資料項目間的關係。

顏色可以傳遞資料項目間的差異和對應。在差異方面,顏色用來表示資料特徵上的相等(或不相等),對應上則可表現離散資料的類別、層級或連續資料的值大小。

互動模式則有參數的(parametric)、視景轉換(view transformation)、過濾(filtering)、揀選細節(details-on-demand)、變形(distortion)。參數的互動是指變更位置、形狀和顏色等參數。視景轉換以縮放、旋轉、翻轉等,對視景提供物理性接觸。過濾透過顏色或形狀等方式,視覺性地選取一部份的資料。揀選細節是指資料的詳細資訊可以被隨時檢索出,並呈現在場景上。變形提供投射的視覺化結果,讓不同的觀點可以同時被觀察與定義。

These efforts are generically known as (data) Visualization, which provides faster and user-friendlier mechanisms for data analysis, because the user draws on his/her comprehension immediately as graphical information comes up to his/her vision.

Several classification schemes have been proposed for visualization techniques, each one focusing on some aspect of the visualization process.

However, many questions remain unanswered.
What are the building blocks of a visual exploration scene?
How interaction mechanisms relate to these facts?

In this work, we discuss these issues and analytically find answers to them based on the very mechanisms of the visualization techniques and on visual perception theory.

In this paper we discuss the subjective nature of visualization by proposing a discrete model that can better explain how visualization scenes are composed and formed, and how their constituent parts contribute to visual comprehension.

It (Keim [13]) maps visualization techniques within a three dimensional space defined by the following discrete axes: the data type to be visualized (one, two, multi-dimensional, text/web, hierarchies/graphs and algorithm/software), the visualization technique (standard 2D/3D, geometrical, iconic, dense pixel and stacked), and the interaction/distortion technique applied (standard, projection, filtering, zoom, distortion and link & brush).

This taxonomy is suitable to quickly reference and categorize visualization techniques, but it is not adequate to explain their mechanisms.

A simpler taxonomy was earlier presented by Schneiderman [27]. It delineates a pair wise system based on a set of data types to be explored, and on a set of exploratory tasks to be carried out by the analyst. This taxonomy, known as task (overview, zoom, filter, details-on-demand, relate, history and extract) by data type (one, two, three, multidimensional, tree and network) taxonomy, was pioneer in analytically delineating visualization techniques.

Another interesting classification is presented by Chi [4], a quite analytical approach, which details visualization techniques through various properties related to a specific visualization model. The taxonomy embraces data, abstraction, transformation and mapping tasks, presentation and interaction.

Tory and Möoller [29] define Scientific Visualization and Information Visualization, respectively, as continuous ([one, two, three, multi-dimensional] versus [scalar, vector, tense, multi-variate]) and discrete (two, three, multidimensional and graph & tree) classes, according to the intuitive perception of their visual modeling.

Wiss and Carr [34] describe a cognitive based taxonomy that considers attention, abstraction and (interaction) affordance in order to discus 3-D techniques.

Visualization can be understood as data represented visually. That is, it takes advantage of spatialization to allow data to be visually/spatially perceived and it relies on visual stimuli to represent data items or data attributes/ characteristics.

Spatialization of data refers to its transformation from a raw format that is difficult to interpret into a visible spatial format.

In fact, Rohrer et al. [25] state that visualizing the non-visual requires mapping the abstract into a physical form, and Rhyne et al. [23] differentiate Scientific visualization and Information visualization based on whether the spatialization mechanism is given or chosen, respectively.

Semiotic theory is the study of signs and how they convey meaning. According to semiotic theory, the visual process is comprised of two phases, the parallel extraction of low-level properties (called pre-attentive processing) followed by a sequential goal-oriented slower phase.

Pre-attentive processing plays a crucial role in promoting visualization’s major gain, that is, improved and faster data comprehension [30].

The work described by Ware [32] identifies the categories of visual features that are pre-attentively processed. Position (2D position, stereoscopic depth, convex/concave shading), Shape (line orientation, length, width and line collinearity, size, curvature, spatial grouping, added marks, numerosity) and Color (hue, saturation) are considered and, according to Pylyshyn et al [21], specialized areas of the brain exist to process each of them (Figure 2).

Visualizing data demands a maximization of just noticeable differences. To satisfy this need, visualizations rely on pre-attentive stimuli - characteristics inherent to visual/spatial entities. Therefore, the data must first be mapped to the spatial domain (spatialized) in order to be pre-attentively perceived.

In this section we identify a set of procedures for data spatialization: Structure exposition, Projection, Patterned positioning and Reproduction.

Structure exposition: data can embed intrinsic structures, such as hierarchies or relationship networks (graph-like), that embody a considerable part of the data significance.

This class comprises visualization techniques that rely on methods to adjust data presentation so that the underlying data structure can be visually perceived. Examples are the TreeMap technique [28], illustrated in Figure 3(a), and force directed graph layouts [8], such as the one illustrated
in Figure 3(b);

Patterned: this is the simplest positioning procedure, with the set of individual data items arranged sequentially (ordered or not) according to one or more directions, linear, circular or according to specific patterns.

Patterned techniques tend to fully populate the projection area and sometimes are referred to as dense pixel displays. Examples include Pixel Bar Charts [15], showed in Figure 4(a), pie charts (circular disposition), depicted in 4(b) and pixel oriented techniques in general [12].

Projection: stands for a data display modeled by the representation of functional variables. That is, the position of a data item is defined by either a well-known or an implicit mathematical function.

In a projection, the information given is magnitude and not order, as in a patterned spatialization. Examples are Parallel Coordinates (one projection per data dimension), Star Coordinates [11] and conventional graph plots, as illustrated in Figure 5;

Reproduction: data positioning is known beforehand and is determined by the spatialization of the system from where data were collected, as exemplified in Figures 6(a) and 6(b).

In reproduction, the data inherits positioning from its original source.

Position is the primary component for pre-attention perception in visualization scenes and it is strictly related to the spatialization process. So, while spatialization is the cornerstone for enabling visual data analysis (as it maps data to the visual/spatial domain), it also dictates the mechanism for pre-attentive positional perception.

Thus, positional pre-attention occurs in the form of Arrangement, Correspondence and Referential, explained in the following paragraphs. These classes derive, respectively, from spatializations Structure Exposition, Patterned and Projection/ Reproduction.

Structure Exposition → Arrangement: specific arrangements can depict structure, hierarchy or some other global property. Without an explicit referential, information is perceived locally through individual inter-positioning of elements and/or globally through a scene overview. For instance,
TreeMap (Figure 3(a)) presents the hierarchy of the data items, and a graph layout (Figure 3(b)) presents network information.

Patterned → Correspondence: the position of an item, either discrete or continuous, determines its corresponding characteristic without demanding a reference. For example, see Figure 10(b) where each of the four line positions maps one data attribute. Other examples are Parallel Coordinates and Table Lens [22], techniques that define an horizontal sequence for placing data attributes;

Projection → Referential: this is the most obvious relation between spatialization and positional pre-attention. Projections have a supporting function whose intervals define referential scales suited to analogical comprehension.

Reproduction → Referential: the position of an element, discrete or continuous, is given relative to an explicit reference, such as a geographical map (Figure 6(b)), a meaningful shape (Figure 7(a)) or a set of axes (Figure 7(b));

In particular, the Shape stimulus embraces the largest number of possibilities to express information: Differentiation, Correspondence, Meaning and/or Relationship.\

Differentiation: the shape displayed discriminates the items for further interpretation, as in Figures 8(a), 9(a) and 10(a);

Correspondence: discrete (Figure 8(a)) or continuous (Figure 8(b)), each noticeable shape corresponds to one informative feature. Proportion (variable sizing) is the most used variation for this practice;

Meaning: the shape displayed carries meaning, such as an arrow, a face or a complex shape (e.g. text), whose comprehension may depend on user’s knowledge and previous experience, as depicted in Figures 7(b) and 8(b);

Relationship: shapes, such as lines, contours or surfaces, denote the relationship between a set of data items, e.g., in Parallel Coordinates, 3D plots and paths in general, illustrated in Figures 5(a) and 7(a).

Color conveys information by Differentiation and/or Correspondence of data items:

Differentiation: colors have no specific data correspondence, they just depict equality (or inequality) of some data characteristic, as it may be observed in the visualizations depicted in Figures 9(a) and 9(b). The coloring of the items, either discrete or continuous, is data dependent or user input dependent;

Correspondence: discrete or continuous, as observed in Figures 7(a) and (b). In the discrete case each noticeable color maps one informative feature, usually a class, a level, a stratum or some predefined correspondence. In the continuous case, the variation of tones maps a set of continuous data values.

Therefore, we must clarify the role of interaction techniques in the visualization scene. We define two conditions for identifying an interaction technique:

1. An interaction technique must enable a user to define/redefine the visualization by modifying the characteristics of pre-attentive stimuli;
2. An interaction technique, with appropriate adaptations, must be applicable to any visualization technique.

The first condition arises from the direct assumption that interaction techniques alter the state of a computational application. In the case of a visualization scene, its basic components (the pre-attentive stimuli) must be altered.

The second condition arises from the need of having interaction techniques that are well defined, which directs us towards generality. An interaction technique, then, must be applicable to any visualization technique, even if not efficiently.

We identify the following interaction paradigms:
Parametric: the visualization is redefined, visually (e.g., scrollbar) or textually (e.g., type-in), by modifying position, shape or color parameters.
View transformation: this interaction adds physical touch to the visualization scene, whose shape (size) and position can be changed through scale, rotation, translation and/or zoom, not necessarily all of them, as in the FastmapDB tool [2];
Filtering: a user is allowed to visually select a subset of items that, through pre-attentive factors such as color (brushing) and shape (selection contour), will be promptly differentiated for user perception.
Details-on-demand: detailed information about the data that generated a particular visual entity can be retrieved at any moment and presented in the scene.
Distortion: allows visualizations to be projected so that different perspectives (positions) can be observed and defined simultaneously.

In proposing this taxonomy we focused on generalizing the rationale of how visualization scenes are engendered and how they are presented to our cognitive system.

Such a general characterization results in a taxonomy that does not rely on specific details on how techniques operate. Rather, it considers their fundamental constituent parts: how they perform spatialization and how they employ pre-attentive stimuli to convey meaning.

Existing taxonomies categorize techniques based on diverse and detailed information on how techniques perform a visual mapping. This diversity and detailing (refer to Section 2) include, e.g., axes arrangement (“stacked techniques”), specific representational patterns (“iconic and pixel-oriented techniques”), predisposition of representativeness (“network and tree techniques”), dimensionality (“2D/3D techniques”) and interaction (“static/dynamic techniques”). Although such approaches can suitably describe the set of available techniques, they lack analytical power because the core constituents of the techniques are diffused within the taxonomical structure.

Our approach results in an extensible taxonomy that can accommodate new techniques as, in fact, any technique will rely on common foundational basis.

2014年4月17日 星期四

Keim, D. A. (2001). Visual exploration of large data sets. Communications of the ACM, 44(8), 38-44.

Keim, D. A. (2001). Visual exploration of large data sets. Communications of the ACM, 44(8), 38-44.

information visualization

視覺資料探索(visual data exploration)將人類的知覺能力運用在大量資料探索過程裡,減少過程中所需要的認知能力,實際上也就是將資料以某種視覺形式呈現,讓資料分析師可以獲得其中蘊涵的洞悉(insight),做出結論,並且與其互動。視覺資料探索運用的時機包括對資料的認識有限以及對於探索的目的模糊等。此外,除可可以讓使用者直接處理資料,相較於自動化的資料探勘技術,視覺資料探索具有可以容易地處理高度不同類的(imhomogeneous)及有雜訊的資料、直覺、不需要了解複雜的數學或統計學演算法與參數等優點。

視覺資料探索的過程大抵上遵循所謂的資訊搜尋箴言(information seeking mantra) [11] 的三步驟,概觀全體 (overview)、放大與過濾(zoom and filter)、選取與觀看細節 (details-on-demand)。相關的技術可以從三個標準進行歸類。
1. 被視覺化的資料類型(the data type to be visualized):一維(如時間資料)、二維(如地圖)、多維(如關連式資料表)、文件與超文件、階層與圖式資料、演算法與軟體。
2. 如何在螢幕上安排資料以及如何處理資料的多維度(multiple dimensions)等視覺化技術(the visualization technique)。
3. 使視覺化產生動態改變及將多個獨立視覺化聯繫與合併的互動(interaction)技術與在深入的同時保留資料全體概觀的扭曲變形(distortion)技術。

視覺資料探索可以根據它們對特定資料特性的適合性進行評估與比較。任務特性則包括叢集(clustering)、分類(classification)、關連(associations)、多變量熱點(multivariate hot spots)等,視覺化的特性包括視覺重疊(visual overlap)和學習曲線(learning curve),希望能夠提供有限的視覺重疊、快速學習和良好的回收。



Visual data exploration seeks to integrate humans in the data exploration process, applying their perceptual abilities to the large data sets now available. The basic idea is to present the data in some visual form, allowing data analysts to gain insight into it and draw conclusions, as well as interact with it.

The visual representation of the data reduces the cognitive work needed to perform certain tasks.

Visual data exploration is especially useful when little is known about the data and the exploration goals are vague.

In addition to granting the user direct involvement, visual data exploration involves several main advantages over the automatic data mining techniques in statistics and machine learning:
• Deals more easily with highly inhomogeneous and noisy data;
• Is intuitive; and
• Requires no understanding of complex mathematical or statistical algorithms or parameters.

A visual representation provides a much higher degree of confidence in the findings of the exploration than a numerical or textual representation of the findings.

Visual data exploration, also known as the “information seeking mantra” [11], usually follows
a three-step process: overview, zoom and filter, and details-on-demand.

These techniques are classified using three criteria: the data to be visualized, the technique itself, and the interaction and distortion method (see Figure 1).

The classification begins with the data type to be visualized [11], including whether it is:
• One-dimensional (such as temporal data, as in Figure 2);
• Two-dimensional data (such as geographical maps, as in Figure 3);
• Multidimensional data (such as relational tables, as in Figure 4);
• Text and hypertext (such as news articles and Web documents);
• Hierarchies and graphs (such as telephone calls and Web sites, as in Figure 5); and
• Algorithms and software (such as debugging operations).

The visualization technique fits into one or more of the following categories, as identified in Figure 1:
• Standard 2D/3D displays using standard 2D or 3D visualization techniques (such as x-y plots and
landscapes) for visualizing the data.
• Geometrically transformed displays using geometric transformations and projections to produce useful visualizations.
• Icon-based displays that visualize each data item as an icon (such as stick figures) and the dimension values as features of the icons.
• Dense pixel displays that visualize each dimension value as a color pixel and group the pixels belonging to each dimension into an adjacent area [6].
• Stacked displays that visualize the data partitioned hierarchically.

The techniques associated with each of these categories differ in how they arrange the data on the
screen (such as 2D display or semantic arrangement) and how they deal with multiple dimensions
in case of multidimensional data (such as multiple windows, icon features, and hierarchy).

Interaction techniques, which allow users to interact directly with a visualization, include filtering, zooming, and linking, thus allowing the data analyst to make dynamic changes of a visualization according to the exploration objectives; they also make it possible to relate and combine multiple independent visualizations.

Interactive distortion techniques support the data exploration process by preserving an overview of the data during drill-down operations. Basically, they show portions of the data with a high level of detail and other portions with a lower level of detail.

Visualization techniques and visual data exploration systems can be evaluated and compared with respect to their suitability for certain data characteristics (such as data types, number of dimensions,
number of data items, and category). Task characteristics include clustering, classification, associations, and multivariate hot spots; visualization characteristics include visual overlap and learning curve.

Desirable visualization characteristics for any technique include limited visual overlap, fast learning, and good recall.

Undesirable visualization characteristics include occlusions and line crossings that might appear to the user/viewer as an artifact limiting the usefulness of the visualization technique.

2014年4月10日 星期四

Fekete, J. D., Van Wijk, J. J., Stasko, J. T., & North, C. (2008). The value of information visualization. In Information Visualization: Human-Centered Issues and Perspectives (pp. 1-18). Springer Berlin Heidelberg.

Fekete, J. D., Van Wijk, J. J., Stasko, J. T., & North, C. (2008). The value of information visualization. In Information Visualization: Human-Centered Issues and Perspectives (pp. 1-18). Springer Berlin Heidelberg.

Information Visualization

Card, Mackinlay, and Shneiderman [2]將視覺化(visualization)定義為「為了增強認知,利用電腦支援、互動的資料視覺表現」。他們並指出歸納視覺能增強認知的方式在於
-- 增加可運用的記憶與處理資源
-- 減少資訊的蒐尋
-- 增強樣式的辨認
-- 產生知覺推理(perceptual inference)運作
-- 使用知覺注意機制進行監控
-- 將資訊以可處理的媒介進行編碼

根據Ware [26],前注意處理理論(preattentive processing theory)和格式塔理論(Gestalt theory)是兩種主要的可以解釋視覺如何有效地知覺特徵和形狀的心理學理論。前注意處理理論(preattentive processing theory)解釋能夠有效處理的視覺特徵,資訊視覺化便是根據前注意處理理論選擇資料呈現的視覺編碼,藉以使感興趣的視覺查詢(visual queries)在前注意處理完成。格式塔理論則是提供描述視覺系統使了解圖像的重要原理,包括接近性(proximity)、相似性(similarity)、連續性(continuity)、對稱性(symmetry)、封閉性(closure)以及相對大小(relative size)等。

本研究認為資訊視覺化的最佳應用為極大資訊空間的探索,當人們尚未知道問題何在或是想提出更好、更有意義的問題時,提供檢視資料來進行了解,產生新的發現或者洞悉資料。Lin [8] 認為這種的瀏覽對資料有良好的相關結構,但同時使用者不熟悉資料集合的內容也僅有限地系統的組織方式,他們對相關資訊需求的描述有困難,對資訊的辨認比描述容易,並偏好以較少的認知負荷進行探索等情形下有幫助。因此,資訊視覺化發揮功用的過程如下:使用者提出一個他們感興趣的問題,將資料以正確的表示方式呈現,讓使用者了解這個表示方式,回答問題並且引發許多預期外的發現與問題。

資訊視覺化的價值雖然可以用使用這項技術的計畫成功來判斷,但由於視覺化往往不是這些成功的唯一方法。本研究則提出以知識增加的價值與需要的成本之間的差來計算資訊視覺化的效益。以數學的方式表示如下:
F = nm(W( ΔK) − Cs − kCe) − Ci − nCu.
其中n代表使用這種視覺化方法的使用者人數,m則是他們的平均使用次數,k則是每次探索需要的步驟數目。而每位使用者每次經由視覺化能獲得的知識價值以W( ΔK))表示,每次需要的前置成本為Cs,反覆進行探索時每一步驟花費的成本則需要Ce,並且這個視覺化技術的研究開發成本和每位使用者選擇與獲得這項技術分別為Ci和Cu。根據上面的公式,視覺化技術要獲得最大的效益需要使用的使用者愈多,然後經常使用來獲取高價值的知識,並且在時間與硬體、軟體與精力上的花費盡量少。

They (Card, Mackinlay, and Shneiderman) describe visualization as “the use of computer-supported, interactive visual representations of data to amplify cognition.” [2]

InfoVis systems are best applied for exploratory tasks, ones that involve browsing a large information space. Frequently, the person using the InfoVis system may not have a specific goal or question in mind. Instead, the person simply may be examining the data to learn more about it, to make new discoveries, or to gain insight about it. The exploratory process itself may influence the questions and tasks that arise.

InfoVis systems, on the other hand, appear to be most useful when a person simply does not know what questions to ask about the data or when the person wants to ask better, more meaningful questions. InfoVis systems help people to rapidly narrow in from a large space and find parts of the data to study more carefully.

Lin [8] describes a number of conditions in which browsing is useful:
– When there is a good underlying structure so that items close to one another can be inferred to be similar
– When users are unfamiliar with a collection’s contents
– When users have limited understanding of how a system is organized and prefer a less cognitively loaded method of exploration
– When users have difficulty verbalizing the underlying information need
– When information is easier to recognize than describe

Information Visualization is about developing insights from collected data, not about understanding a specific domain.

Information Visualization is still an inductive method in the sense that it is meant at generating new insights and ideas that are the seeds of theories, but it does it by using human perception as a very fast filter: if vision perceives some pattern, there might be a pattern in the data that reveals a structure.

Following that definition, the authors (Card, Mackinlay, and Shneiderman [2]) listed a number of key ways that the visuals can amplify cognition:
– Increasing memory and processing resources available
– Reducing search for information
– Enhancing the recognition of patterns
– Enabling perceptual inference operations
– Using perceptual attention mechanisms for monitoring
– Encoding info in a manipulable medium

According to Ware [26], there are two main psychological theories that explain how vision can be used effectively to perceive features and shapes. At the low level, Preattentive processing theory [19] explains what visual features can be effectively processed. At a higher cognitive level, the Gestalt theory [6] describes some principles used by our brain to understand an image.

Preattentive processing theory explains that some visual features can be perceived very rapidly and accurately by our low-level visual system. ... Information visualization relies on this theory to choose the visual encoding used to display data to allow the most interesting visual queries to be done preattentively.

Gestalt theory explains important principles followed by the visual system when it tries to understand an image. According to Ware [26], it is based on the following principles:
Proximity Things that are close together are perceptually grouped together;
Similarity Similar elements tend to be grouped together;
Continuity Visual elements that are smoothly connected or continuous tend to be grouped;
Symmetry Two symmetrically arranged visual elements are more likely to be perceived as a whole;
Closure A closed contour tends to be seen as an object;
Relative Size Smaller components of a pattern tend to be perceived as objects whereas large ones as a background.

Still, the process of explaining how InfoVis works remains the same: ask a question that interests people, show the right representation, let the audience understand the representation, answer the question and realize how many more unexpected findings and questions arise.

One effective line of argumentation about the value of InfoVis is through reporting the success of projects that used InfoVis techniques. These stories exist but have not been advertised in general scientific publications until recently [16,12,9]. One problem with trying to report on the success of a project is that visualization is rarely the only method used to reach the success.

Visualization can be considered as a technology, a collection of methods, techniques, and tools developed and applied to satisfy a need. Hence, standard technological measures apply: Visualization has to be effective and efficient.

The profit of visualization is defined as the difference between the value of the increase in knowledge and the costs made to obtain this insight.

A schematic model is considered: One visualization method V is used by n users to visualize a data set m times each, where each session takes k explorative steps.
The value of an increase in knowledge (or insight) has to be judged by the user. Users can be satisfied intrinsically by new knowledge, as an enrichment of their understanding of the world. A more pragmatic and operational point of view is to consider if the new knowledge influences decisions, leads to actions, and, hopefully, improves the quality of these. The overall gain now is nm(W( ΔK)), where W( ΔK)) represents the value of the increase in knowledge.

Concerning the costs for the use of (a specific) visualization V , these can be split into various factors. Initial research and development costs Ci have to be made; a user has to make initial costs Cu, because he has to spend time to select and acquire V , and understand how to use it; per session initial costs Cs have to be made, such as conversion of the data; and finally during a session a user makes costs Ce, because he has to spend time to watch and understand the visualization, and interactively explore the data set. The overall profit now is

F = nm(W( ΔK) − Cs − kCe) − Ci − nCu.

In other words, this leads to the obvious insight that a great visualization method is used by many people, who use it routinely to obtain highly valuable knowledge, while having to spend little time and money on hardware, software, and effort.

The costs Ce that have to be made to understand visualizations depend on the prior experience of the users as well as the complexity of the imagery shown.

The costs Cs per session and Cu per user can be reduced by tight integration with applications.

The initial costs Ci for new InfoVis methods and techniques roughly fall into two categories: Research and Development.
Research costs can be high, because it is often hard to improve on the state of the art, and because many experiments (ranging from the development of prototypes to user experiments) are needed. On the other hand, when problems are addressed with many potential usages, these costs are still quite limited.
Development costs can also be high. It takes time and effort to produce software that is stable and useful under all conditions, and that is tightly integrated with its context, but here also one has to take advantage of the large potential market. Development and availability of suitable middleware, for instance as libraries or plug-ins that can easily customized for the problem at hand is an obvious route here.


2014年3月28日 星期五

Yi, J. S., Kang, Y. A., Stasko, J. T., & Jacko, J. A. (2008, April). Understanding and characterizing insights: how do people gain insights using information visualization?. In Proceedings of the 2008 Workshop on BEyond time and errors: novel evaLuation methods for Information Visualization (p. 4). ACM.

Yi, J. S., Kang, Y. A., Stasko, J. T., & Jacko, J. A. (2008, April). Understanding and characterizing insights: how do people gain insights using information visualization?. In Proceedings of the 2008 Workshop on BEyond time and errors: novel evaLuation methods for Information Visualization (p. 4). ACM.

information visualization

洞悉經常被認為是資訊視覺化的結果,但獲得洞悉的過程仍無法被了解,目前資訊視覺化的文獻將獲得洞悉的過程分為四類(提供全貌、調整、偵測樣式、比對心智模式),這些過程提供了一些瞭解。Saraiya et al. [23]認為洞悉是觀察資料後的一個發現,可分為以下的四種情形:全貌(overview)、樣式(patterns)、群組(groups)以及細節(details)。North [16] 提出複雜(complex)、深入(deep)、質性(qualitative)、不可預期(unexpected)和相關(relevant)是洞悉的特性。本研究援引 Pirolli and Card [17]認為洞察為意義建構(sense-making)的一個部分,而意義建構如Klein et al. (p.71) [11]所定義之為有動機且持續的努力來了解人事間的連結,藉以預測它們進行的軌道(trajectories)與有效地行動。基於此一定義,意義建構有如下的現象:首先,意義建構的過程是周而復始的循環(cyclic and iterative);其次,意義建構不只是發現的過程,更是創造的過程;最後,意義建構是回溯的(retrospective),人們通常先建構一個架構,然後回溯地蒐集相關資料,並將它放到架構上,如果蒐集的資料與架構相符合,這個架構便獲得確定,否則人們會感到困惑,會捨棄、更正或取代原先的架構,來解釋新的資訊。Klein et al. [11]用Figure 1來說明上述的特性。


本研究以文獻分析探討資訊視覺化的研究者認為人們如何透過資訊視覺化獲得洞悉?的看法,歸納出四種不同但交融(intertwined)的方式:
1) 概觀(overview):概觀描述人們對整個資料集合進行全盤了解的過程,能夠提供人們掌握已知與未知的事務、發現值得進一步探索的區域以及此一資料集合可以供新知的範圍。
2) 調整(adjust):調整描述人們在探索資料集合時的過程調整抽象層級(the level of abstraction)與選取範圍(the range of selection)的過程。在探索大量資料時,運用選取功能可以進行過濾;群集(grouping)的功能可以聚集、簡化、組織和標示相關資料,使得大量資料能夠進行管理。
3) 偵測樣式(detect pattern):偵測樣式表示發現資料集合內的特定分布、趨勢、頻率、離群(outliers)與結構。
4) 比對心智模式(match mental model):資訊視覺化能夠降低了解時的認知負荷(cognitive load),增強現有事物的再認知(recognition),並且將呈現的視覺資訊與實際的知識連結起來。
循環交互地運用這些過程能夠提供對於資料集合的洞悉,可以對應到上述的意義建構理論。但由於這些過程得到的洞悉都相當抽象而高層次,資料本身的特質、使用者的興趣與背景知識對獲取洞悉有相當大的影響。

We found that: 1) Insights are often regarded as end results of using InfoVis and the procedures to gain insight have been largely veiled;
2) Four largely distinctive processes of gaining insight (Provide Overview, Adjust, Detect Pattern, and Match Mental Model) have been discussed in the InfoVis literature;
and 3) These different processes provide some hints to understand the procedures in which insight can be gained from InfoVis.

Saraiya et al. [23] have conducted insight-based evaluation studies in the domain of biology (analyzing biological pathways and micro-array data), and they define insight as “an individual observation about the data by the participant, a unit of discovery” (p. 444).

They (Saraiya et al. [23]) further group insights found in the context of bioinformatics into four different categories: overview (overall distributions of gene expression), patterns (identification or comparison across data attributes), groups (identification or comparison of groups of genes), and details (focused information about specific genes). (p. 445).

North [16] describes characteristics of insight as follows (p. 6):
Complex. Insight is complex, involving all or large amounts of the given data in a synergistic way, not simply individual data values.
Deep. Insight builds up over time, accumulating and building on itself to create depth. Insight often generates further questions and, hence, further insight.
Qualitative. Insight is not exact, can be uncertain and subjective, and can have multiple levels of resolution.
Unexpected. Insight is often unpredictable, serendipitous, and creative.
Relevant. Insight is deeply embedded in the data domain, connecting the data to existing domain knowledge and giving it relevant meaning. It goes beyond dry data analysis, to relevant domain impact.

Sensemaking clearly begets insights as shown in the model of sensemaking (Information -> Scheme -> Insight -> Product) proposed by Pirolli and Card [17].

Sensemaking simply can be defined as “making sense of things” or, drawing from Klein et al. (p.71) [11], more comprehensively described as “a motivated, continuous effort to understand connections (which can be among people, places, and events) in order to anticipate their trajectories and act effectively.”

First, sensemaking procedures are cyclic and iterative. Russell et al. [22] describe the sensemaking procedure using Learning Loop Complex theory, which consists of 1) search for representations (generation loop), 2) instantiate representations (data coverage loop), 3) shift representations. As the name of their theory illustrates, the process of sensemaking is iterative in collecting data and re-generating a representation or scheme.

Second, sensemaking is not only a discovery procedure, but also a creation procedure. Weick [32], an organizational theorist, emphasized this notion of creation by providing a very clear distinction between interpretation and sensemaking. Weick argues that interpretation is a component of sensemaking, and states “the act of interpreting implies that something is there, a text in the world, waiting to be discovered or approximated. Sensemaking, however, is less about discovery than it is about invention” (p.13).

Third, sensemaking is retrospective. Literature in various contexts (e.g., [7,32]) supports that people often do not make sense of things after collecting information first. Instead, people often construct a framework first and retrospectively collect the relevant information and place it into the framework. If the collected information fits well with the framework, the framework is confirmed. However, if it does not, people become puzzled, and the framework could be discarded, updated, or replaced to explainthe new information.

These characteristics are well summarized by the Data/Frame Theory of sensemaking of Klein et al. [12].

Insight is not only an end result or simple discovery of hidden truth, but also an intermediate state in the iterative and cyclic procedure of sensemaking and invention. Insight could be a framework that needs to be created first in a person’s mind to draw the boundary of a problem and collect and understand information.

Hence, we conducted an extensive InfoVis literature review, specifically focusing on the following question: “How do people gain insight through InfoVis?” We reviewed 4 books, 2 book chapters, and 34 papers published in major venues in InfoVis. In order to consider the bigger picture, we initially focused on books and articles that survey the benefits of InfoVis, and latter reviewed case studies and evaluation studies to find some procedural aspects of insight.

In this section, we introduce four largely distinctive processes through which people gain insight while using an InfoVis system. ... The processes we have identified are: 1) Provide Overview, 2) Adjust, 3) Detect Pattern, and 4) Match Mental Model.

Provide Overview characterizes processes through which a person comes to understand the big picture of a dataset of interest. Even though observing an overview may not directly help a person gain insight, it appears to play an important role by helping people make sense of and find which areas they need to investigate more, thereby promoting further exploration of the dataset.

That is, Provide Overview allows people to grasp what they know and do not know, which areas are available for further investigation, and to what extent they could gain new knowledge from the dataset.

Adjust refers to a process through which people explore a dataset by adjusting the level of abstraction and/or the range of selection. Being able to flexibly change perspective on the dataset allows people to make sense of various aspects and test different hypotheses they have generated.

Selecting the range of a dataset to display by using a filtering interaction technique is a way to help explore a large amount of data.

Grouping is also an effective way to explore data by abstracting huge datasets into more manageable pieces. ... Through the process of grouping and aggregating, relevant information is gathered, simplified, organized, and labeled.

Detect Pattern means to find specific distributions, trends, frequencies, outliers, or structure in the dataset. ... A pattern itself could be an insight and further a person can cast a new questions and hypotheses by understanding patterns.

One of the benefits of InfoVis is that a visual representation of data can decrease the gap between the data and user’s mental model of it, thereby reducing cognitive load in understanding, amplifying human recognition of familiar presences, and linking the presented visual information with real-world knowledge.

Instead, these four different processes are intertwined and often used together to generate insights. For example, Provide Overview, as mentioned previously, often precedes further Adjust, and Adjust and Detect Pattern are often used together to gain deeper insight. This aspect clearly mirrors the cyclic and iterative characteristic of sensemaking.

More specifically, some of insights found in the literature are somewhat abstract and higher level, so that linking them with any of the categories is questionable.

Additionally, one of most important factors to help users gain insight might be the degree of users’ engagement into the dataset. ... The nature of data and users’ interests and background knowledge can also heavily affect the insight gaining procedure.

Usability seems to be another important aspect to promote the insight gaining process.

Clutter and occlusion are also examples of barriers for insight acquisition that need to be addressed. Sometimes too much data on a limited screen results in visual clutter and occlusion, which in turn diminishes the possibility of uncovering patterns and trends. Consequently, researchers have sought to reduce clutters in various ways [6].

2014年3月27日 星期四

Chang, R., Ziemkiewicz, C., Green, T. M., & Ribarsky, W. (2009). Defining insight for visual analytics. Computer Graphics and Applications, IEEE, 29(2), 14-17

Chang, R., Ziemkiewicz, C., Green, T. M., & Ribarsky, W. (2009). Defining insight for visual analytics. Computer Graphics and Applications, IEEE, 29(2), 14-17.

information visualization

許多研究都指出資訊視覺化的目的是提供洞悉(insight),例如 Card, Mackinlay and Shneiderman [1]與Thomas and Cook [2]。然而目前大多數對於洞悉的定義卻是莫衷一是,例如North便有兩種不同但相關的看法。North [3] 認為洞悉的特徵包含複雜(complex)、深入(deep)、質性(qualitative)、不可預期(unexpected)以及相關(relevant),這一類的看法與認知科學上的突發性洞悉(spontaneous insight)相近,將洞悉認為是靈光一現的剎那(a moment of enlightenment),也就是問題從不知道如何解決忽然轉移到知道如何解決的一個過程,並且這類的看法需要注意的是這類的問題解決的過程並非依循尋常的方式,而是在僵局下,透過微弱的語意網絡,忽然激發較不清楚地相關資訊,所產生的典範轉移(paradigm shift)。但North與其同事[5]也曾提出另一種關於洞悉的看法,他們將洞悉定義為參與者對於資料的一種個別觀察(an individual observation)以及是一個發現的單位(a unit of discovery),這種看法可以說是將洞悉視為是一種知識的進步(an advance of knowledge)或是一片段的資訊(a piece of information),這種看法認為視覺化能幫助知識的建構,例如Yi et al. [6]以意義建構理論(sense-making theories)為基礎將視覺化能產生的洞悉分為四個不同但彼此交疊的過程:提供全貌(provide overview)、調整(adjust)、偵測樣式(detect patterns)以及比對心智模式(match mental model)。

由於自發性的洞悉來自於語意知識(semantic knowledge)不可預期的重組(reconfiguration),要對一個問題產生自發性的洞悉必須具有相關知識。在另一方面,自發性的洞悉所引起的典範轉移(paradigm shifts)能夠讓人對於問題的瞭解產生新的結構與關係。因此,本研究認為這兩種看法在習得知識的循環上彼此相互支持,以Figure 2來表示。在僅有有限知識的一開始(0 to k1),使用者並無法產生自發性的洞悉。當知識逐漸增加後(k1 to k2),能夠產生自發性洞悉的可能性增加。最後(k2 to k3),愈多的知識能夠產生自發性洞悉的可能性愈增加,但趨勢逐漸減緩。所以,應提供一個環境讓兩種看法的洞悉都能發生。


Many have argued that providing insight is the main goal of information visualization. Stuart Card, Jock Mackinlay, and Ben Shneiderman declare that “the purpose of visualization is insight,” [1] while Jim Thomas and Kris Cook propose in Illuminating the Path that the purpose of visual analytics is to enable and discover insight [2].

For example, Chris North categorizes insight to be “complex, deep, qualitative, unexpected, and relevant,” [3] which overlaps with the neurological definition.

However, North and his colleagues also define insight as “an individual observation about the data by the participant, a unit of discovery,” [5] which does not bear any clear relation to the strict aha moment of cognitive science. Instead, it implies a focus on knowledge-building not found in the cognitive definition.

We suggest that what the visualization community defines as insight actually has two parallel meanings: a term equivalent to the cognitive science definition of insight as a moment of enlightenment, and a broader term to mean an advance in knowledge or a piece of information.

The cognitive science community has used the term insight “to name the process by which a problem solver suddenly moves from a state of not knowing how to solve a problem to a state of knowing how to solve it.” [8]

In this tradition, spontaneous insight is a type of problem solving and  differs from normal problem solving in several key ways.
First, spontaneous insight doesn't appear to  be facilitated by gradual learning heuristics such as bottom-up inductive reasoning.  In fact, researchers have observed that focused effort on normal problem solving often inhibits spontaneous insight. Spontaneous insight usually occurs when a person is in a relaxed state [9] (such as when taking a shower in the morning).
Second, whereas gradual problem solving requires no special inducement other than presenting someone with a problem, what precipitates spontaneous insight is still being discussed.  One commonly held theory is that spontaneous insight often occurs when a person tries to solve the problem in a habitual way, fails, momentarily becomes frustrated (perhaps owing to incorrect assumptions or some other cognitive fi xedness), mentally reorganizes the pieces of the puzzle (perhaps by breaking through a failed thought paradigm), and “suddenly” sees the solution. [8]
Finally, in normal problem solving the path taken to the solution is conscious and logically clear to the problem solver; however, participants who experience a spontaneous insight often can’t describe the thought process that led to it, [10] indicating that this insight occurs subconsciously and isn't a process that can be directly controlled, manipulated, or repeated.

This indicates that normal problem solving involves a narrow but continuous focus on information highly relevant to the problem at hand. ... . This suggests that spontaneous insight occurs through sudden activation of less clearly relevant information through weak semantic networks, which corresponds to a participant’s paradigm shift following an impasse.

These findings suggest that spontaneous insight is qualitatively different from everyday problem solving. It involves a unique pattern of neural activity that corresponds with the unique sensation of 
the “aha” moment that participants report.

Recently, Yi and his colleagues provided a comprehensive survey on information visualization literature that considered insight as a goal or a measurement [6]. On the basis of sense-making theories, they concluded that four distinct but intertwined processes in visualization can lead to insight: provide overview, adjust, detect patterns, and match mental model.

In the visualization community, researchers often talk about discovering insight, gaining insight, and providing insight. This implies that insight is a kind of substance, and is similar to the way knowledge and information are discussed.

In the cognitive science community, researchers more often discuss experiencing insight, having an insight, or a moment of insight. In this context, insight is an event.

On the basis of the cognitive definition of insight, this statement restricts visualization into considering only a specific mode of problem solving that produces results that, although measurable, aren't easy to track.

On the other hand, considering insight only as knowledge or information limits visualization’s potential to structured knowledge building and information display.

If spontaneous insight comes from the unexpected reconfiguration of semantic knowledge, [10] then relevant knowledge about a problem must be necessary for spontaneous insight to arise. ... Conversely, the major paradigm shifts associated with spontaneous insight can create new structures and relationships in a user’s understanding of a problem, which can then serve as the schematic structures needed for generating future knowledge-building insights.

Together, the two types of insight support each other in a loop that allows human learning to be both flexible and scalable.

As Figure 2 shows, when the user has only a limited amount of knowledge (0 to k1), spontaneous insight won’t likely occur.
As the amount of knowledge increases (k1 to k2), the probability of spontaneous insight increases sharply.
Finally, after a certain point (k2 to k3), further increase of knowledge increases the probability in only a limited fashion until it’s asymptotically close to a spontaneous insight occurring.
On the other hand, a reduction in the probability of gaining a spontaneous insight undoubtedly occurs, at least for a while, if the user is distracted from this freer knowledge association.
But whatever model is chosen, our main point is that spontaneous and knowledge-building insights should be considered distinct because the best approaches to gain one or the other are different.

For spontaneous insight, we can evaluate exploratory, “prequery” approaches that keep one “in the cognitive zone” or “in the flow,” and quantitatively identify when a spontaneous insight occurs through an EEG or fMRI.

For knowledge-building insight, we can evaluate detailed knowledge-gathering methods and look to appropriate user studies to measure how much knowledge a user gains.

Using these combined approaches, we can not only more accurately determine visualization tool’s effectiveness, but also provide cognitive scientists with more complex problem-solving artifacts (they have few available) and shed light onto how to promote the two types of insight through visualization tools to solve real-world problems.