顯示具有 information visualization 標籤的文章。 顯示所有文章
顯示具有 information visualization 標籤的文章。 顯示所有文章

2017年11月30日 星期四

Chen, S., Arsenault, C., Gingras, Y., & Larivière, V. (2015). Exploring the interdisciplinary evolution of a discipline: the case of Biochemistry and Molecular Biology. Scientometrics, 102(2), 1307-1323.

Chen, S., Arsenault, C., Gingras, Y., & Larivière, V. (2015). Exploring the interdisciplinary evolution of a discipline: the case of Biochemistry and Molecular Biology. Scientometrics102(2), 1307-1323.

本研究從跨學科性(interdisciplinarity)的變化、核心學科的確定、學科的出現和潛在的學科檢測幾個方面,探討生物化學與分子生物學(Biochemistry and Molecular Biology, BMB)一百年間的科際整合演化,並利用科學地圖和河流圖(StreamGraph)做為檢視科際整合演化的視覺化工具

Porter和Rafols(2009)研究六個研究領域的三十年間跨學科性的演變,研究結果顯示,在這段時間內有顯著的變化,特別是每篇文章引用的學科數量和參考文獻的數量,以及每篇文章的合著者數量。不過,Rao-Stirling指數只是略有上升,Porter和Rafols認為這是由於論文的引用還是主要是在相鄰的學科領域內。Larivie`re and Gingras (2014)的研究發現,在1945和1975間跨學科研究下降後,已經逐漸上升。Levitt et al. (2011)研究SSCI (Social Science Citation Index)的各主題類別的跨學科性,則是發現1980和1990之間下降後,2000年已恢復1980年的水準。

在文獻計量學,跨學科性的測量通常是根據跨學科的引文關係。例如,Porter和Chubin(1985)使用「外部類別引用」(Citations Outside Category, COC)的指標衡量文章的跨學科性程度,這個指標定義為來自不同學科的引用或參考文獻的百分比。這種方法被應用到多個研究,例如Rinia等人(2002a)測量所有科學領域的跨學科性,Morillo et al. (2001) 測量化學領域的跨學科性。Adams等人(2007)除了不同學科的引用參考文獻的比例,也使用了引用的學科數量和Shannon多樣性指數 (Shannon Diversity Index)。Carley and Porter (2012)利用Rao-Stirling多樣性指數(Rao-Stirling diversity index),透過引用文獻,探索知識整合。Levitt和Thelwall(2008)則是使用WoS和Scopus的主題分類,以論文發表期刊的多重分類測量跨學科性

Rinia等人(2002b,244)將跨學科性定義為一組研究人員在他們的“主要”學科之外發表的文章的百分比。Qiu (1992)使用作者的組織隸屬關係探究跨學科合作。Abramo等人(2012)以義大利為案例研究,以研究人員的學科作為類別指定的方法,利用不同學科的合作者之間的合作確認跨學科性。Le Pair(1980)和Sugimoto等(2011)則是利用作者最高學歷學科探索跨學科Le Pair(1980)將跨學科合視為科學家在他們的生涯中從一門學科向另一門學科的轉移。Sugimoto等(2011)利用80年(1930-2009)的博士論文,從學術譜系描述了圖書情報學跨學科性變化程度。

本研究利用WoS資料庫上主題類別BMB下從1910年起所有的文章,共 1,539,526篇,以及其文件類型為期刊論文的參考文獻,共40,855,852 筆。參考文獻的期刊以美國國家科學基金會(National Science Foundation, NSF)的主題分類作為學科指定的系統,共143個分類,每種期刊主要被指定到一個分類,但研究上僅考慮超過250筆參考文獻的類別,並且以參考文獻所占比例較多的類別做為核心學科,此外也預測有潛力的學科。計算上使用兼顧引用文獻分布的學科數量和集中性(concentration)以及學科間相似性Rao-Stirling指數 (Rafols and Meyer, 2010)來分析COC。本研究使用相對開放性(relative openness)指標(Lee等人, 2009)來衡量引用學科的影響,並通過查看這些學科的影響的變化來探索科際整合的演變。

在視覺化部分,本研究使用河流圖表現核心學科長期變化,河流的寬度變化表示學科影響力的增加或減少。將按照時間順序出現學科可視化是另一種說明BMB科際整合的方法,本研究選擇引用次數等於或超過250次的學科探究BMB的新興學科。另外,本研究以VOSveiwer產生科學地圖,圖上的每一個單位為NSF主題分類的143個學科,其距離與大小由10年間(2003~2012年)的學科共被引矩陣計算,並依據NSF的14個主要學科分類著上顏色。




本研究的結果顯示,在100年間每年BMB引用的學科從1成長到93,跨越12個主要學科分類,只有Humanities和Arts兩個主要學科分類未曾被引用主要學科分類中較重要的有Clinical Medicine、Biomedical Research和Biology,另外,Physics和Chemistry也相當重要。從引用學科增長的數量,可見BMB在100年期間愈來越具有跨學科性。


比較各學科被BMB引用的比例(ri)以及它們的論文占所有發表論文的比例(pi),本研究確認出15個核心學科:



核心學科的首次引用大約都出現在1950年之前,較晚出現的核心學科大都是新興的學科。並且在科學地圖上,核心學科通常被映射在BMB的附近,也就是說與BMB較近的學科在早期便被BMB引用。

雖然BMB的引用主要來自本身(45.32%),但引用的比例明顯在下降中,1910年為74.0%,但2012年只有32.1%。除了BMB本身以外,一般生物醫學研究(General Biomedical Research)是另一個最多引用的學科。然而隨時間改變,不同的時期有不同重要的科際整合學科。在早期,生理學(Physiology)、藥學(Pharmacology)和免疫學(Immunology)是最重要的學科,然而近期這些學科的重要性減少,取而代之的是細胞生物學(Cellular Biology)、細胞學和組織學(Cytology and Histology)、遺傳學(Genetics and Heredity)與癌症學(Cancer)等多種學科,以及一般生物醫學研究。










在剛開始的時期,BMB主要引用本身學科的論文,然後是化學(Chemistry, 1924)、臨床醫學(Clinical Medicine,
1937)和生物學(Biology)。這三個學科也是在整個NSF分類系統所形成的科學地圖上與BMB最接近的學科。隨著BMB的發展,距離較遠的學科也逐漸加入。根據BMB的跨學科發展順序,可以分為四個時期:第一時期為1910到1960年的50年間,這個時期共有來自生物醫學、臨床醫學、化學和生物學的18個學科,在這個時期內核心學科大多已經加入。第二時期為1961~1981年的20年,這時期新增了34個學科,這時期除了生物醫學和臨床醫學的引用大量增加之外,特別值得一提的是來自物理學的參考文獻增加,顯示BMB的跨學科範圍開始從它鄰近的學科向外擴大。第三時期則是從1982到2002年的20年,這個時期新增了27個學科,跨學科的範圍更擴大到工程與科技(Engineering and Technology)、心理學、地球與空間和數學。第四時期則是從2003年開始,共有16個新學科,特別一提的是這個時期加入的圖書資訊學(Library and Information Science),其原因是這個時期開始有研究利用書目計量學方法對BMB的研究進行評估。從四個時期的科學地圖也可以發現,BMB引用的學科愈來愈多元,而且從一開始使用的學科都是在科學地圖上較接近的學科,愈到後期愈多來自距離較遠的學科。




以下則是最近十年增長最快的前五學科,這些學科可以視為是具有與BMB具有跨學科潛力的學科。




在100年間,Rao-Stirling指標從約0.3變為約0.6,成長約2倍,與Porter and Rafols (2009)的研究相比,同一時間區間(1975~2005年),Porter and Rafols (2009)研究的6個學科,Rao-Stirling指標並沒有明顯增加,但BMB則約增加0.32倍。

本研究證實跨學科主要從相鄰領域向認知較遠的領域演變,並且BMB研究者引用其他學科文獻日益增加,而雖然引用較遠領域的比例較小,但正在顯著增加,因此是BMB較有跨學科潛力的學科



2017年8月23日 星期三

Viegas, F. B., Wattenberg, M., & Feinberg, J. (2009). Participatory visualization with wordle. IEEE transactions on visualization and computer graphics, 15(6).

Viegas, F. B., Wattenberg, M., & Feinberg, J. (2009). Participatory visualization with wordle. IEEE transactions on visualization and computer graphics15(6).

Wordles (http://www.wordle.net/)是一個產生文字雲(word cloud)的線上服務,與tag clouds相似,兩者都是採用字型大小來展示文章中詞語的出現頻率,然而相較於tag clouds按照字母順序排列,而且僅有較單調的(同一色調)的字型,wordles還包括多種變化的顏色與配置,產生的圖形較令人印相深刻,使用者容易陶醉在各種可能的顏色、印刷樣式(typography)與組合(composition)之中。藉由研究網路上的Wordles使用情形和數千個受訪者的調查,本論文探討wordles受歡迎的原因:能展現使用者的創造力。
本論文發現Wordles經常使用在主流媒體(mainstream media)、個人使用(personal usage)和在教育上使用(usage in education)等三種場合。將Wordles在主流媒體上的用途,例如用來比較政治人物的修辭以及提取網路上的使用者生成內容(user-generated content)。許多使用者會將Wordles運用在對他們有意義的文本上,從部落格(blogs)、Twiiter上的內容,到詩詞或歌詞,都是可以輸入Wordles的文本。並且使用者不僅將產生的作品放在網路上與其他人分享,甚至放在T恤、卡片上。老師與學生在課堂上的使用也是Wordles常見的用途之一。
根據4300多位受訪者的調查結果,Wordles受歡迎的原因主要有二:一為Wordles可以讓使用者作為一種創作的工具;其次Wordles產生的設計結果具有吸引人的外觀。許多使用者會充分利用不同的字型、顏色和配置,來創作自己的作品。然而調查的結果卻也指出有明顯比例的使用者不了解字型大小代表的意義。與同樣文本產生的tag clouds相比,有70%的使用者較喜愛wordles,歸納受試者偏好Wordles的原因包括:1. 受到豐富色彩所引發的情感影響(emotional impact);2. 同時兼具吸引人的注意和較長記憶保持的視覺效果(attention-keeping visuals);3. 使用者沉醉在Wordles違反許多可讀性的非線性(non-linearity)閱讀上。87%的受訪者想要讓Wordles產生的視覺化效果更接近文本內容而嘗試不同的配置方式或指定字型和顏色。受訪者運用Wordles創作文字雲的主要原因包括:樂趣(fun)、創造力(creativity)、教育目的(educational purpposes)和做為禮物(gift giving)。
論文引用Jenkins提出的參與式文化(participatory culture)[5]來分析wordles現象。參與式文化中包含「對於藝術表現(artistic expression)和公民參與(civic engagement)有較小的障礙、產生與分享個人創作有較強的支持、以及某種形式的非正式指導(informal mentoring)」等要素。在這種匯流的文化(culture of convergence)中,消費者被鼓勵尋求新的資訊,並且與分散的媒體內容產生連結。參與式文化將使用者展現在新媒體系統上的工作與玩樂一起涵蓋起來,因此綜合以上的討論,Wordles可以視為是一種參與式文化根據本論文的研究結果,作者認為視覺化工具不應只是依賴於科學風格的資料分析(scientific-style data analysis),而且也是一種創作工具(authoring tool),因此使用者(特別是非專業的使用者)不僅只是使用這些工具觀察作品,而是能應用協助他們創作或是將資料混合而重新創作(remixing)的工具。而從觀看工具轉變成創作工具的過程主要有三個關鍵:1. 使用者可以自行選擇喜愛的顏色、配置方式,乃至於分享方式,如此一來,每一個Wordles的作品都是獨一無二的,使用者會對他們的作品更有擁有感(ownership)。2. 每一個作品除了在Wordles網站上有一個專屬的網址,也能夠匯出成為PDF檔案,並且作品都在Creative Commons的法律框架下受到保護,這個機制讓使用者將他們的Wordles作品視為財產,鼓勵進一步的實驗與創作。3. 使用者能將對他們有意義的文本,運用Wordles進行創作,並且分享。


Wordles are close relatives of tag clouds, encoding word frequency information via font size.

We suggest that a key message of the Wordle phenomenon is that scientific-style data analysis is not the only raison d'être of visualization tools.

Our results suggest that Wordle usage may be viewed as a component of “participatory culture” (in the sense of Jenkins [5]): a cultural system in which viewers are also producers and remixers, and where visualization serves as much as an authoring tool as a method of analysis.

Wordle qualifies as a casual infovis system [6] , since it may be used by non-expert users to depict personally meaningful information. However, we will argue that the usage scenarios that we see with Wordle go beyond the definition of casual visualization in several ways, especially in the many cases where it seems to function as a remixing or authoring tool.

In fact, as we’ll argue, the response to Wordle hinges so strongly on the notion of user creativity that it may be fruitfully viewed in the framework of Jenkins [5], who discusses “participatory culture.” In his definition, this is a culture that includes “low barriers to artistic expression and civic engagement, strong support for creating and sharing one’s creations, and some type of informal mentorship.”

Usage in the mainstream media

A favourite way to use Wordle was to compare the rhetoric of the two major parties. For example, both the Washington Post and the Boston Globe featured a side-by-side comparison of the Democratic and Republican candidates’ blogs (fig 2), with the takeaway being that “Obama” was, by far the most prominent word on McCain’s campaign blog.

Often Wordles were used to distil user-generated content from the web. The Guardian wordled the outlook for 2009 as defined by Twitter users, for instance, while a WIRED magazine piece entitled “Mourning the Internet Famous: Randy Pausch’s Distributed Funeral” used a Wordle to illustrate the top 20 most used words in comments from Tributes.com on the professor’s death.

Personal Usage

Many users created wordles of personally meaningful data. Sources range from online media (blogs, Twitter feeds, etc.) to poetry and song lyrics.

One common pattern is for a group of people to find each other online and create a set of related wordles.

Part of the excitement around Wordle seems to lie in the ability to take visualizations beyond the Web. Because the site lets users export a high-resolution version of their creations, savvy users can put wordles onto T-shirts, cards, and other physical objects.

Usage in education

A striking number of sites gave extensive tutorials on how to use the tool in the classroom. (Perhaps one might expect this from an audience of teachers!)

Two main themes emerged from our results: the importance of design and that the Wordle site works like a creation tool.

Design and visual appeal were overwhelmingly cited as a reason for users’ interest in Wordle. Not only were wordles attractive, the design inspired users to engage with the visualization in creative ways.

When asked which representation was more effective, 70% of participants felt Wordles were more effective compared to 11% who found tag clouds outperformed Wordles (19% of respondents felt both representations did equally well).

Participants’ preference for Wordle can be broadly grouped into three main categories: emotional impact, attention-keeping visuals, non-linearity.

87% of respondents used Wordle’s customization capabilities— either by trying different layouts or specifying combinations of font and color. ... The frequency of edits speaks to users’ interest in experimenting with typographical arrangements that fit their needs.

The fact that users make ample use of font, color, and layout choices points to a second theme: a feeling of creativity in using Wordle.

 At least within our survey sample, wordles were not passively consumed: 76% of our respondents said they had personally produced one.
Most respondents used texts they were already familiar with, with the majority wordling their own writing (table 1).

Finally, the survey included a freeform question where respondents could elaborate on why they had created word clouds. Four main themes emerged: fun, creativity, educational purposes, and gift giving.

Overall, younger people (under 20) tended to know the least about how each different dimension worked, with females doing worse than male users. One set of answers in particular stands out: word size. 35% of young males and 49% of young females did not understand the meaning of word size. Older females too (above 30) did not do so well: 31% did not understand what word size meant.

Terminology aside, it’s worth pointing out two cognitive processes supported by Wordle that are not directly related to statistical analysis or insight: learning and memory.

The ability of wordles to assist learning and memory seems directly related to their aesthetic qualities.

We argue that the feeling of creativity is central to the experience of using Wordle. Even the examples where Wordle aids learning and memory include elements of creation.

The activity surrounding Wordle seems to fit Jenkins’ definition of participatory culture.

The barriers to entry are low, and we see both self expression and, in the political Wordles, civic engagement.

The system has technical and legal infrastructure for sharing creations.

Finally, there is indeed informal mentorship.

The range of uses, from frivolous to serious, is characteristic of the participatory culture arc, with more whimsical uses leading to more sophisticated analysis—the element of “fun” attracts novice users of a system [5], and helps them learn how it works.

If it is true that Wordle has made the transition in people’s minds from a viewing tool to an authoring tool, one might ask how it has done that. We suggest there are three ingredients: user choice, artifact portability, and remixing power.

User choice: Choice is inherent in the Wordle experience. Users can experiment with everything from color and layout to ways of sharing. The implications of such flexibility are twofold: each Wordle has the potential to look unique, and users are more likely to take ownership of their work. As our survey indicates, this ownership is a key element in how people relate to Wordle: by giving them choice, Wordle becomes an artifact of a participatory medium.

Artifact portability: An authoring tool must be able to produce a lasting artifact. Several features of Wordle ensure that users’ creations can persist. Not only can people create a web page with a persistent URL, as is typical of many online visualization sites, but it is easy to export to PDF. Beyond these technical capabilities, the legal framework of the Creative Commons license helps people distribute their creations. This infrastructure empowers users to think of Wordles as their property, inspiring further experimentation and creativity.

Remixing power: Everyone has text they care about, whether emails, love letters, or speeches made by a hated politician. Text is almost never just “data.” Pointing Wordle to the latest cultural meme, be it a speech by the president or the stimulus bill, and then sharing it, proves a quick and easy way of engaging in communal and civic meaning-making.

Our study revealed two main themes behind Wordle’s broad uptake: the importance of design and the fact that the Wordle site works as a creation tool.

At the same time, our survey revealed some potentially problematic aspects of the Wordle experience. A significant number of people do not understand the information encoding in Wordle. Our survey indicated strong age and gender differences in how wordles were interpreted, which suggests natural directions for future research. In addition to testing these findings in a lab setting, one might extend this investigation to how well the average person understands other very simple charts and graphs.


2016年1月12日 星期二

Engelhardt, Y. (2007). Syntactic structures in graphics. Computational Visualistics and Picture Morphology, 5, 23-35.

Engelhardt, Y. (2007). Syntactic structures in graphics. Computational Visualistics and Picture Morphology5, 23-35.

圖形的目的是用來說明某種的資訊,使得原本無法看見的訊息能被看見 (visualizing the nonvisual)。依據過去的文獻,本研究建議將圖形的建構(building blocks)分為三個部分: a) 顯示的圖形物件 (graphic objects)、b) 配置這些物件使其具有意義來說明訊息的圖形空間 (graphic spaces)以及c) 這些物件的圖形性質 (graphic properties)。在此,圖形物件為一個遞迴的結構 (recursive structure),也就是:在一個圖形空間上配置的一組圖形物件能夠共同形成一個更高階的圖形物件。

本研究也提出圖形物件的語法類型 (syntactic categories),用來解釋物件彼此間可以允許的空間關係並且區分圖形的基本構成 (basic constituents)。所有的圖形都是建立在不同語法類型的圖形物件之組合的可能性上,不同語法類型的圖形物件在圖形呈現 (graphic representation)上有不同的行為,此一限制產生它們有不同的空間定位。所有圖形物件的語法類型可分為兩大群組:1) 依附於圖形空間位置上的物件;2) 依附於其他物件上的物件。前者包括節點 (node)、線標示 (line locator)、面標示 (surface locator)以及格標記 (grid marker);後者則有標籤 (label)、連結線 (connector)、比例區段 (proportional segment)和框架 (frame)。以地圖為例,節點 、線標示、面標示可分別代表地圖上城市、河流與湖泊或國家的標示,格標記則用於表示地圖上的經緯度,標籤則是標明節點代表的城市名稱。各種圖形物件的類型、依附類型與例子,如下表:




根據語言的語法學概念,圖形的語法學 (the syntactics of graphics)研究不同語法類型的圖形物件之間的關係,研究圖形物件與圖形空間之間以規則與限制為基礎的關係,也研究圖形物件如何組合成複合的圖形物件以及複合的圖形物件如何能以更簡單的圖形物件分析。圖形可分為實體場景和物件的影像以及抽象的圖形,前者如照片與地圖等呈現實體的空間,後者則如家族樹 (family trees)、統計圖表 (statistical charts) 等呈現概念性的空間。照片與地圖等利用影像中的空間配置呈現出真實的空間配置;家族樹和圓餅圖 (pie charts) 則以影像中的空間配置呈現非空間性的資訊。然而,實體空間的呈現並不一定表達出被呈現物件的真實座標比例(the true coordinate proportions),許多圖形也同時由實體和概念性的空間組合而成,例如在地圖上以高度呈現國家的人口密度分布。下表是若干圖形空間的類型與代表。


Building upon the existing literature, we are suggesting to regard the building blocks of all graphics as falling into three main categories: a) the graphic objects that are shown (e.g., a dot, a pictogram, an arrow), b) the meaningful graphic spaces into which these objects are arranged (e.g., a geographic coordinate system, a timeline), and c) the graphic properties of these objects (e.g., their colors, their sizes).

We suggest that graphic objects come in different syntactic categories, such as nodes, labels, frames, links, etc. Such syntactic categories of graphic objects can explain the permissible spatial relationships between objects in a graphic representation.

In addition, syntactic categories provide a criterion for distinguishing meaningful basic constituents of graphics.

It is about images that can be regarded as ‘visualizing the nonvisual’ in an attempt to clarify information of some sort. Such images are often collectively referred to as “graphics”.

In 1914, Willard Brinton writes in his book Graphic methods for presenting facts that “The principles for a grammar of graphic presentation are so simple that a remarkably small number of rules would be sufficient to give a universal language”.

In 1967, Jacques Bertin publishes his classic Sémiologie graphique, in which he analyses the “language” of graphic representations and the “visual variables of the image”.

In 1976, linguist Ann Harleman Stewart examines the properties of diagrams and claims that “Like any language, graphic representation has a vocabulary and a grammar”.

In 1984, Clive Richards proposes a “grammatically-based analysis” of diagrams in his Ph.D. thesis Diagrammatics.

In 1986, Jock Mackinlay suggests that “graphical presentations are actually sentences of graphical languages that have precise syntactic and semantic definitions”. In Mackinlay’s approach, “the syntax of a graphical language is defined to be a set of well-formed graphical sentences”.

In 1987, Fred Lakin publishes his paper “Visual grammars for visual languages”, in which he describes his approach to the “spatial parsing” of graphics, which he defines as “the process of recovering the underlying syntactic structure of a visual communication object from its spatial arrangement”.

Kress and van Leeuwen publish their book Reading images: the grammar of visual design (1996). Unfortunately, it is difficult to extract a systematic approach to a syntactic analysis of graphics from their book.

A paper titled “The visual grammar of information graphics” (1996) by Engelhardt et al., suggests “syntactic categories of visual components”.

Robert Horn, in his book Visual Language (1998), proposes a morphology and a syntax of visual language based partly on the work of Jacques Bertin and on the Gestalt principles of perception.

In his book The grammar of graphics (1999), Leland Wilkinson describes an approach to graphics that is related to object-oriented design in computer science. However, he uses grammatical terminology “metaphorically”, and not in a linguistic sense.

Colin Ware (2000) writes about the “perceptual syntax of diagrams”, describing “the grammar of node-link diagrams” and “the grammar of maps”.

Engelhardt
, in his Ph.D. thesis The language of graphics (2002) provides a detailed proposal for the analysis of syntactic structure, which he applies to a broad spectrum of graphic representations.

We propose a notion of graphic objects that will allow for recursive structures: Any graphic representation – and any meaningful visible component of a graphic representation – may be referred to as a graphic object. This means that graphic objects can be distinguished at various levels of a graphic representation. For example, a map or a chart in its entirety is a graphic object. In addition, the various symbols or components that are positioned within that map or chart are graphic objects as well.

A bottom-up description of this principle was given above: a set of graphic objects can be arranged into a graphic space, together forming a single graphic object at a higher level. This “nesting” or “embedding” (Engelhardt 2002) of graphic structures can be referred to as “recursive composition” (Card 2003).

In technical terms, a meaningful graphic space could be defined as a graphic space that involves an interpretation function from spatial positions to one or more domains of information values.

In graphics, not only the possible constituents themselves (graphic objects), and the diverse possible ways of arranging these constituents (in meaningful graphic spaces), but also the possible visual appearances of these constituents (graphic properties such as size, color), could be considered as being part of the graphic “vocabulary”. In this sense we can say that the building blocks of graphics fall into three main categories: graphic objects, meaningful graphic spaces, and graphic properties.



To make a more general statement, we claim that all graphics are based on the possibility of combining graphic constituents (graphic objects) of different syntactic categories (Engelhardt et al. 1996, Engelhardt 2002, 2006).

Graphic objects of different syntactic categories “behave” differently in a graphic representation. The constraints that govern their spatial positioning are different.




All syntactic categories of graphic objects can be divided into two main groups: 1) objects that are attached to locations in graphic space (e.g., node, line locator, surface locator, grid marker are all attached to locations in graphic space), and 2) objects that are attached to other objects (label, connector, proportional segment, frame are all attached to other objects).

Richards (1984) believes that “there seems to be little profit in using such items as an individual dot or line as a unit of analysis. If we are going to use linguistics as a model, then what is needed for present purposes is not the pictorial equivalent of a phoneme or morpheme but something closer to a noun phrase”.

The basic graphic objects in a particular graphic representation are those that can be regarded as functioning in some syntactic category within that particular graphic representation (e.g., as a label, as a node, as a connector, as a proportional segment, etc.).

The distinction between syntactics, semantics, and pragmatics was introduced by Charles Morris (1938, 1946). Morris conceives of syntactics as the investigation of the relationships between signs, of the ways in which complex signs can be constructed from simple ones, as well as the ways in which complex signs can be analyzed into more simple ones (Morris 1946/1971).

The syntactics of graphics investigates the relationships between graphic objects of different syntactic categories. It investigates the rule- and constraint-based relationships between graphic objects (of different syntactic categories) and graphic spaces.

And syntactics investigates how graphic objects can be combined into composite graphic objects, and how composite graphic objects can be analyzed into more simple ones.

Looking at the broad spectrum of graphics we can say that images of physical scenes and objects, such as pictures and maps, represent physical spaces, while many abstract graphics, such as family trees and statistical charts, represent conceptual spaces (Engelhardt 1999, 2002).

In other words, pictures and maps use spatial arrangement in the image to represent spatial arrangement in the world, while family trees and pie charts use spatial arrangement in the image to represent non-spatial information.

Representations of physical spaces do, by the way, not always have to express the true co-ordinate proportions of the represented objects.

Many graphics combine physical and conceptual spaces.

As an example of a true hybrid space (Engelhardt 1999, 2002), think of a three-dimensional landscape drawing of a country in which the drawn “mountains” do not represent physical mountains, but – for example - population density, peaking in the cities and flat in the countryside. In this case, the horizontal plane represents the physical space of the country’s geography, while the vertical dimension represents the conceptual space of population density.




We claim that all types of graphic representation of information can be analyzed in terms of their composition from graphic spaces of different sorts.

We have tried to show that specifying such a visual language means a) specifying the syntactic categories of its graphic objects, plus b) specifying the graphic space in which these graphic objects are positioned, plus c) specifying the visual coding rules that determine the graphic properties of these graphic objects (see table 1).

The syntactic structure of a graphic representation is determined by the rules of attachment for each of the involved syntactic categories (see table 2) and by the structure of the meaningful graphic space that is involved (see table 3).

With this analysis we have attempted to demonstrate that Morris’ original notion of syntactics applies well to the structure of graphics.

2016年1月3日 星期日

Wanner, F., Stoffel, A., Jäckle, D., Kwon, B. C., Weiler, A., Keim, D. A., ... & Pfister, H. (2014). State-of-the-art report of visual analysis for event detection in text data streams. In Computer Graphics Forum (Vol. 33, No. 3).

Wanner, F., Stoffel, A., Jäckle, D., Kwon, B. C., Weiler, A., Keim, D. A., ... & Pfister, H. (2014). State-of-the-art report of visual analysis for event detection in text data streams. In Computer Graphics Forum (Vol. 33, No. 3).

近年來從文本串流中偵測事件已成為熱門的研究領域,然而由於對事件的概念沒有妥善的定義以及文本資料的種類繁多,對資料分析與視覺化是一個重大的挑戰,因此能夠處理特定事件類型與多樣性文字來源的視覺分析工具的建立準則十分缺乏。在本研究中,將事件視為是從文本資料中抽取出的對使用者有價值的非預期而獨特的樣式(unexpected and unique patterns),並且建議由新聞標準(news criteria)或新聞價值(news values)[GR65]來界定事件的價值。

在從文本資料串流利用視覺分析進行事件偵測的研究中,資料的來源從有限而書寫良好的新聞文章到社交媒體上由使用者書寫、快速產生甚至有時沒結構的文字資料;而分析任務則可依其目的分為新事件偵測(new event detection)、事件追蹤(event tracking)、事件摘要(event summarization)以及事件關聯(event associations) (Dou et al., DWRZ12)。Becker [Bec11]對於社交媒體上的事件偵測研究,將事件依據3個面向區分:1) 計畫內 (planned) vs. 無計畫 (unplanned)、 2) 趨勢 (trending) vs 非趨勢 (non-trending)、 3) 外源 (exogenous) vs. 內源 (endogenous)。

本研究以Figure 1上的流程圖表示事件偵測與探索的處理過程,首先在輸入文件資料的前處理 。在前處理之後的方法,則可分為兩類:一類首先應用自動化方法偵測資料裡的事件,然後再利用這些資訊做為視覺分析(visual analysis)的介面;另一類則是直接對前處理後的結果進行視覺化,不進行事件的自動化分析。






以下分別說明文本資料來源、文本處理方法等技術分析的面向。

文字資料來源包括:新聞、電子郵件、部落格、RSS feed、微網誌 (microblogging) 訊息、論壇(forum)上的發文(post)、客服表單、影像與視頻串流上附註的文字。目前有大半的研究是針對微網誌資料,例如Twitter,而除了文字以外,微網誌上的地理位置和作者等後設資料也是許多研究會加以利用的。

在文句偵測(sentence detection)、(tokenizing)、詞幹化(stemming)和(lemmatizing)等文本處理方法之後,進行較深入的詞類標示(part-of-speech tagging)、語法剖析(syntactic parsing)、文句中詞語關係的類型剖析 (Typed-dependency parsing)、相互指涉解析(coreference resolution)、專有名詞辦認 (named entity recognition)、極性抽取 (polarity extraction)、歧義消除 (word sense disambiguation)等。自2000到2011年,33篇事件偵測的相關論文只有17篇利用文本處理方法,但在2012年後,這個情形改變了,18篇論文中便有14篇論文使用文本處理方法,主要是詞類標示和極性抽取。

事件偵測的自動化方法中常用的技術可分為1)群集為基礎、2)以分類為基礎、3)以統計為基礎、4)以預測為基礎、5)本體論為基礎、6)模式探勘(pattern mining)、7)資料串流重複特徵的模型、8)規則式等類型。
以群集為基礎的方法將文件依據內容的不同特性分群,當群組改變時便是事件產生。
分類則是根據事件建立分類器,當文件的分類結果為事件相關,便視為是事件產生。
統計方法中,相關分析類的方法測量文件集合在詞語或詞語與時間上的相關性改變來偵測事件,另一類則是從稀少或獨特的詞語出現來發現事件。
預測為基礎的方法根據過去的歷史預測接下來文件的出現情形。
本體論(ontologies)為基礎的方法適合單一領域的事件偵測,以全自動或半自動方法產生特定本體,當偵測到活躍概念(activated concepts)中的改變時便是可能的事件。
模式探勘(pattern mining)利用A-priori 演算法抽取文件串流上的常見連續模式。
資料串流重複特徵的模型可以用來偵測事件,當一個串流明顯偏離它預期的特徵時便是偵測到一個事件。
規則為基礎的方法以人工編寫的規則偵測事件,例如根據詞語或詞頻為規則來偵測特定的事件。

在視覺化呈現上,以時間為基礎的視覺化呈現佔有大多數,共21篇論文,利用包括河流 (river)、時間線 (timeline)與圓形 (circular)等時間為基礎的呈現凸顯資料在時間上的演變。時間線的呈現會將符號放置在一或多條時間線上,表現資料項目、密度與數量,能夠表現單一或稀少的事件。河流通常用來做為群集演算法結果的視覺化,提供各群集的分布與整體的數量,較著重在高頻率的事件上。以微網誌為資料來源的應用,通常會利用微網誌上附加的地理參考資料,以地圖的方式來呈現。折線圖或長條圖等基本的視覺化呈現方式通常運用來表現事件有關的資料在時間上的數量與頻率。在各種視覺化的應用中,文字資料的呈現通常選用具有意義的關鍵詞。

在支援的分析任務上,包括 1) 提供文件集合的概觀,描述集合內發現的主題,利於進一步的分析。2) 關鍵詞語搜尋以及利用後設資訊過濾資料,在分析或視覺化時減少資料或抽取的事件數量。3) 監測資料來源中事件在時間上的發展。4) 將偵測到的不同事件間的關係視覺化。

質性的評估方法包括案例研究(case study)、使用性評估(usability evaluation)、使用案例(use case)以及軼事評估(anecdotal evaluation),其中使用案例最為盛行。較常見的量化評估方式比較偵測到的事件與真實的資料。


Event detection from text data streams has been a popular research area in the past decade.

However, data analysts and visualization experts often face grand challenges stemming out of the ill-defined concept of event and various kinds of textual data. As a result, we have few guidelines on how to build successful visual analysis tools that can handle specific event types and diverse textual data sources.

Within this paper, events are regarded as unexpected and unique patterns extracted from text data streams, valuable to users.

In particular, data sources evolved from a relatively limited amount of well-written news articles to rapidly generated, user written, and in some cases unstructured textual data from social media services.

Dou et al. [DWRZ12] defined task according to “New Event Detection”, “Event Tracking”, “Event Summarization”, and “Event Associations”, but we expect that tasks can be even more diversified including geographic dimension which were not explored yet.

Another challenge represents the unstructured, diverse textual data. It mandates extensive processing and preparation in order to properly employ it.

Becker [Bec11] shows interesting work about event detection in social media. She divides an event using three dimensions: 1) “planned” vs. “unplanned”; 2) “trending” vs “non-trending”; 3) “exogenous” vs. “endogenous”. The last dimension aims to detect events within the data in a real-life context.

Some examples of visual social media analysis is shown in Schreck and Keim [SK13]. With screenshots of the different visualizations, the authors explain the underlying data, analysis methods, and functionality of various applications in visual social media analysis.

There exists a survey on semantic sensemaking by Bontcheva and Rout [BR12]. Though their focus was on the semantic aspects, a subsection refers to visualization approaches.

Rohrdantz et al. [ROKF11] mention tasks for the “RealTime Visualization of Streaming Text Data”. They call tasks that are relevant in terms of the scope of our paper “monitoring”, “change and trend detection” and “situational awareness”.



In the first step of the pipeline, the documents are prepared for the analysis. In this step the documents are parsed to get the plain texts and standard text preprocessing methods, such as sentence detection, tokenizing, and stemming and lemmatizing are applied. In addition to these standard methods, methods from the computer linguistic field can be used in the preprocessing step to annotate the texts with additional information. For instance, part-of-speech tagging, named entity extraction, or syntactic parsing can be used to identify types of words, persons and places, or structure of sentences.

After the preprocessing step different approaches are used to detect events (see two branches in Processing in Figure 1).
The first group of approaches applies automatic methods to detect patterns in the data. The detected patterns are then used to create a visual analysis interface for the data set, what we call visual analysis. The interaction between visualization and the automatic part shapes a visual analytics approach.
The second group of approaches skips the automatic analysis and directly visualizes the outcome of the preprocessing, what also is only visual analysis because of the lack of interaction possibilities of a certain extent.

Text Data Sources

1. News is a well-known text data source. News captures information of a real world event or happening. It consists of a title, often followed by a short summary and the body containing details about the event. News goes through a professional gatekeeping process which in the end forms the agenda of media.

2. A typical electronic document is email. ... Emails are used for personal conversations, advertisement or business information exchange. They consist of a header and a body. The header contains information about transaction: sender, receiver, timestamp, and other meta data. The body contains the textual content of the email. An email body can be of arbitrary length which is one of its characteristics.

3. Weblogs, shortly named blogs are used for information purposes of a more or less undefined audience. ... A blog can have a specific topic or can be open for various topics.

4. RSS feeds are a standardized format to broadcast short news snippets. They consist of a title and a description. RSS feeds can be used by news agencies, newspapers and blogs. ... The standardized format allows the easy integration into other applications.

5. Recently, microblogging providers are becoming more and more popular. The messages are limited with respect to their length of 140 characters. So-called “hashtags” are used in order to characterize the membership of a tweet to a certain topic. In addition, more meta data is provided, e.g. geolocation, author, place etc.

6. User forums often have hierarchical structure. A message within the forum is a post and is not strictly restricted with respect to its length. Posts which belong to the same topic shape a so called thread. On the other hand several threads often belong to a sub-forum within the main forum. The purpose of a forum is the discussion on specific issues and topics regarding the its main topic.

7. Modern customer-care systems often ask each customer to fill out a feedback form after a purchase. This form (often digital, reachable through the internet) gives the customer the opportunity to provide issues directly to the vendor. The information is a valuable source which allows the seller to react fast and adequately to issues being raised by customers. ... Often these forms are semi-structured, which means they have checkboxes for predefined questions and provide a free text field for further comments.

8. Images and video sequences can be uploaded on sharing sites such as Flickr (https://www.flickr.com/). Users can tag their content with text. These tags and little text snippets typically describe the content in a short manner or express an emotional state being associated with the photo.

Almost half of the papers use microblogging data namely Twitter. It is obvious in Table 1 that in 2010 a shift towards microblogging happened. It is also noticeable that meta data (geolocations, author information) is often used in conjunction with microblogs.

Text Processing Methods

1. Part-of-speech (POS) tagging detects the word type of tokens.

2. Syntactic parsing determines the grammatical structure of sentences. ... Full syntactic parsing uses grammars and build up a complete parse tree for a sentence. ... Shallow parsing creates meaningful chunks and avoids the complexity of full parsing.

3. Typed-dependency parsing determines the type of relations between words in a sentence.

4. Coreference resolution creates connection between referring expression, such as pronouns, and subjects in a text. A correct resolution of referring expressions could improve text mining results, e.g., polarity extraction would benefit from correctly resolved referring expressions.

5. Named entity recognition (NER) detects and labels names of, e.g., persons, locations, events, or dates in texts.

6. Polarity extraction or determines the attitude (positive vs. negative) of the writer about a subject.

7. Word-sense disambiguation techniques use the context of words to determine the correct sense of tokens. ... We only observed one paper using word sense disambiguation.

We confirm that text processing methods are used very sparingly. ... Since 2000 until the end of 2011, only 17 out of 33 papers utilized any of the methods. The 16 papers with no text processing methods solved the event detection tasks with visualization. In the year of 2012, the trend changed dramatically; 14 out of 18 papers have used text processing methods in the papers published since then.

It is also noticeable that part-of-speech tagging and polarity extraction have gained popularity since 2012 as well.

Thus, we believe that many research papers started absorbing more natural language processing techniques to further generate their event metrics.

Automatic Methods for Text Event Detection



1. Clusters are generated for different time windows based different properties in the document, e.g., co-occurrence of terms, frequency in time, or metadata. Events are generated when the set of clusters changes, e.g., a new cluster arise or two existing clusters merge.

2. Users provide a set of example documents and classifiers learn to detect the annotated events. Classifier-based techniques are used in similar cases with rule-based ones, but have the advantage that users do not need to create rules by themselves.

3. Statistical methods such as correlation or detection of outliers and significant difference are used to identify events. Correlation based methods examine collection between terms or between terms and time and detect events by changes in the correlation measures. A different type of statistical methods calculate term-wise deviation from an expected value or use other measures to identify rare or unique occurrence of terms.

4. Prediction-based methods predict the occurrence of following documents based upon past history.

5. Methods based on ontologies [HHSW09] are suitable for event analysis in single domains. Specific ontologies are generated with full- or semi-automatic methods. ... Using this type of methods, events can then be detected from changes in activated concepts.

6. Pattern mining algorithms, such as the A-priori algorithm of Wu and Chen [WC09] applied to text in [WSJ∗ 14], are used to extract common sequential patterns in document streams. Patterns can be found based on documents themselves or time intervals. In both cases, features extracted from documents are then used to define patterns.

7. Models of the recurring characteristics of data steams can be used to detect events. An event is detected when a stream deviates significantly from its expected characteristics.

8. Rule-based approaches detect events with manually created rules. For instance, users specify rules based on terms and/or frequency to detect a particular event.

Visualization of Events in Text Data



In total, 21 papers use a time-oriented visualization (river, timeline, circular) to visualize the evolution of the data over time. Time-oriented visualizations are often combined with additional visualizations to show non-time dependent information.

Maps visualization came up with microblogging data and use mainly geographic references in the meta information of the microblogs for visualization.

A problem for all visualizations is the question how to visually represent text data. This problem is usually solved by selecting meaningful keywords that are either generated by frequency or by another scoring technique such as topic models.

Basic visualizations (e.g. line or bar charts) are mainly used to give an overview of the data set by showing the time dependent relations of events. They are used to visualize the data volumes or frequencies over time, for instance, of detected topics, named entities, or keywords.

Timeline visualization use one or multiple timelines and place glyphs or shapes on these timelines to indicate single data items, densities, or volumes. Timeline visualizations are therefore preferred over river visualizations when single or rare items should be tracked, because a river visualization put the focus on high frequent events.

River metaphors are often used to visualize outcomes of cluster algorithms. Although timeline techniques could be used, rivers provide a space saving overview and give a better visual impression of the distributions of the clusters and the overall amount of data.

 Supported Analysis Tasks

1. Overview visualization give users a summary of the document collection. Common are textual summaries based on frequent terms or topic models that describe the topics found in the collection. These summaries serve as navigation support and are often used as starting point for further analysis.

2. It is also common to provide users with abilities to search for keywords or allow filtering of the data by meta information. Both tasks reduce the number of item or extracted events in the analysis or visualization.

3. Monitoring tasks are the second most frequent tasks supported by the surveyed systems. Users monitoring a data source are interested in the evolution of events in a changing data source. Time-based visualizations (e.g., timeline, river, circular) are often used for monitoring task, because they show the temporal development of events in data sources.

4. In many cases relations between different detected events are visualized. The most frequent shown relations are relation in content, time, and volume. For instance, a river visualization shows time and volume relations between different streams and with additional annotations also relations in content can be shown.

The statistical methods are combined with any type of visualizations.

Interestingly, clustering methods are often visualized by river visualizations.

Exceptionally, topic modeling techniques are not only used with time dependent visualizations but also with other visualizations such as treemaps or geographic visualizations. This pattern appears because topic models are clustering methods that return a ranked list of terms representing single topics, which are often used in visualizations to label data and find names for clusters.

Evaluation


We subdivide qualitative methods into the following categories: case study, usability evaluation, use case, and anecdotal evaluation.

Table 7 accentuates the popular usage of use cases; except for 16 of all considered papers the authors make use of this method. Typically, a use case validates through the description of a fictitious scenario that pinpoints main features whereas a case study involves a domain expert and therefore is more time-consuming [DNKS10,MBB∗ 11].

Anecdotal evaluation describes how the suggested system could be used, but do not provide sufficient evidence to judge the general efficacy of the presented technique.

Usability evaluations involve users performing particular tasks with the given system and asks for comments on usability.

The most prominent quantitative evaluation methods are comparisons of the detected events with a ground truth set. Often event databases are used as ground truth that are enriched by the authors with missing entries.

A different evaluation form of algorithms are comparison with existing algorithms and reporting quality measures. In some cases not the results of the algorithms are evaluated but the performance in the sense of runtime or memory consumption is assessed, which is important for systems working in near real-time scenarios.

We also found only four papers using a user study for evaluation. We expected more papers using user studies, because many systems present novel visualization techniques and user studies can verify the strength and weakness of the application [HHN00,LYK∗ 12,RHD∗ 12].

One thing we noticed was that data sources have dramatically changed from news to social media since 2010. Mainly due to the burst of social media, many research studies used text data streams generated out of Facebook or Twitter.

Some data sources – like for instance discussion forums – are underused than others. Discussion forums are traditional methods to collect opinions from many people, but few research topics investigate data because they are asynchronous and slow to build up in nature. Despite these limitations, they also have a strength: archival history of some topics. Several discussion forums include years of textual conversation between multiple users on a single topic. For instance, this longitudinal conversation can be used to detect certain noticeable shifts in a specific user group’s opinions on political issues over some months or years.

More importantly, visualizations were primarily used as presentation, but had no interaction possible to steer the underlying data processing algorithm in order to further analyze data in a different angle. This limitation can prevent users from providing their insights back into the visualizations.

Especially for news, news criteria (also known as news values) [GR65, HO01] can help find and develop new features for content-based feature detection. They are only mentioned once in our whole bulk of surveyed news analysis research papers [DNKS10].

According to [GR65], news criteria are: frequency, threshold, unambiguity, meaningfulness, consonance, unexpectedness, continuity, composition, reference to elite nations, reference to elite people, reference to persons, and reference to something negative.

In general, news and selection criteria could be merged into one concept we call event values. Event values are a concept including the text data producer’s and user’s perspectives. They could be implemented in the data analysis process by means of new features (feature engineering) and interactive elements, which comes along with the call for more visual analytics functionality.

2015年12月23日 星期三

Gan, Q., Zhu, M., Li, M., Liang, T., Cao, Y., & Zhou, B. (2014). Document visualization: an overview of current research. Wiley Interdisciplinary Reviews: Computational Statistics, 6(1), 19-36.

Gan, Q., Zhu, M., Li, M., Liang, T., Cao, Y., & Zhou, B. (2014). Document visualization: an overview of current research. Wiley Interdisciplinary Reviews: Computational Statistics6(1), 19-36.

文件視覺化(document visualization)是一種資訊視覺化技術,將詞語、文句、文件或它們之間的關係等文字資訊轉換為視覺形式,使得使用者在面臨大量的文件時可以更好的了解文件、減輕他們的心理負荷。文件通常較缺乏結構(minimally structured),但有豐富的特徵(attributes)和後設資料(metadata),因此相較於文本視覺化(text visualization),文件視覺化主要著重在文件以及其包含的特徵和後設資料上。以下是幾種可能的應用:(1) 詞語的頻次與分布; (2) 語意內容與重複 (semantic content and repetition); (3) 區別文件集群的主題; (4) 文件的核心內容;(5) 文件間的相似性;(6) 文件間的連結;(7) 文件內容改變的過程;以及 (8) 社交媒體上的資訊擴散與其他模式以及做為改善文本搜尋的方式。

本研究以視覺化的對象(visualization objects)與任務對蒐集到的文件視覺化技術進行分析,視覺化的對象分為單一文件、文件集合、串流文本訊息以及檢索結果,以下對各種視覺化任務進行說明:

單一文件的視覺化目的在快速了解與吸收核心內容與文本特徵,著重在詞語、片語、語意關係和內容上,分為三種類型:
1.呈現詞語頻次、分布與語彙結構等語彙特徵的語彙為基礎 (Vocabulary-Based)視覺化:
重要的技術有Tag Clouds [6,7]與Wordle [8,9],這類技術利用位置、顏色與大小等方式呈現單一文件中的詞語頻次,近來parallel tag clouds (PTC) [10]、 ManiWordle [11]、 context preserving dynamic word cloud[12]和visualization of internet discussion with extruded word clouds. [13]等許多研究以這類方法為基礎,並加以改善。其他屬於這類但原理不同的方法還有TextArc [14]和DocuBurst [16]。
2. 呈現實體與其間關係的語意結構視覺化:
Semantic Graphs [19]利用Penn Treebank產生的剖析樹(parse tree),產生每一句子內的主語-動詞-受語,在解決代名詞指代(pronominal anaphors)問題,將各實體相連,產生語意圖(semantic graph)。
3.呈現文件內容(Document Content)的特性與關係為基礎的視覺化:
例如WordTree [23]以樹狀結構表現詞語的上下文脈絡,樹的根節點是使用者選取的詞語,每一個分支表示詞語在文件上的上下文,節點大小代表各詞語的頻次。Arc Diagrams [25] 以半圓形的弧連結重複的次序列,用來顯示內容上重複的複雜模式(complex patterns of repetition)。

文件集合的視覺化在於顯現文件群集上的主題、文件間的相似與差異以及內容在時間上的改變,相關技術可分為
1. 文件主題的視覺化:目的在發現特定的主題以及反應各個不同主題之間的關係,著名的研究案例有ThemeScapes [26]和 INSPIRE的 ThemeView以及 The Galaxy [27],分別以地形圖和散佈圖表現文件在主題上的分布情形,地形圖上各「山脈」的高度表示主題的強度。TopicNets則以網路圖的節點與連線呈現文件間在主題上的關係。ThemeRiver [29]和Topic Island [30]著重在文件集合內主題在時間的變化,以ThemeRiver [29]來說,X軸表時間,Y軸上則以不同顏色的「河流」代表各主題,河流的寬度表現主題主題在相關文件上的強度。
2. 文件核心內容的視覺化:目的在提供整個文件集合的概觀,例如Document Cards [35]用來呈現大量的文件集合,每一張卡片上包含文件上重要的詞語和影像,而詞語是由文本探勘(text mining)技術由文件上的文本抽取出來,影像則從文件上抽取或然後加以組合。
3. 版本更動的視覺化:呈現各版本上的差異。例如:History Flow [36] 的設計是用來顯示維基百科上不同版本的文件內容更動情形以及相對應的作者;另外,也有許多針對軟體程式碼發展的視覺化。
4. 文件關係的視覺化:發現在不同文件上實體的連結,實體包括人、地點、日期和組織等等,提供這個功能的視覺化技術如Jigsaw [46] 。其他的技術,如ContexTour [47]和PivotPaths [49] 為視覺化被使用在論文集合的應用、FacetAtlas [48]則將Google Health文件上的病因、症狀、處方和診斷等實體相連,
5. 文件相似性的視覺化:其目的在將相似的文件置於彼此接近的位置,並且能遠離不相似的文件。過去常利用自組織映射圖 (self-organizing map, SOM) [50],將高維度的資料映射到2維平面上呈現,並使得資料間複雜而非線性的關係能夠以距離方式表現,著名的例子有 Lin 的研究[52]和WEBSOM [53]。

在文本視覺化的應用方面,由於近年社交媒體的盛行,即時的串流文本處理的研究大為盛行,研究問題包括主題與詞語的統計分析與表現、主題相關事件的凸顯以及文本訊息本身的視覺化 [54],Christian Rohrdantz [55]進行了串流文本資料的即時視覺化相關研究的回顧,Whisper [57] 的研究可以追蹤社交媒體上的資訊擴散過程。另一個應用是檢索結果的視覺化介面,早期的研究成果如TileBars [58],Sparkler[59]可同時將多個查詢問句的結果以視覺化的方式呈現,RankSpiral [60] 的視覺化呈現重點在比較多個查詢問句或不同搜尋引擎的檢索結果。

各種技術所提供的取用方法、對文件的要求與主要特色可參考下表



關於主要特色說明如下:
1. 擴充性 (Extension):此方法可適用於大量的文件集合。
2. 多功能性 (Versatility):適用於多種的視覺化任務。
3. 互動性 (Interactivity):提供使用者比較直覺的人機介面,讓使用者參與研究與發展過程。
4. 技術 (Techniques)
在文本處理上採用可擴充 (scalable)、高效能 (high-performance) 的演算法;採用協調多視圖 (Coordinated and Multiple Views);即時處理技術。

本研究並且提出文件視覺化有待發展的兩個研究方向,一為文件視覺化的評鑑方法 [62],另一為理論基礎 [63]。


This overview introduces fundamental concepts of and designs for document visualization, a number of representative methods in the field, and challenges as well as promising directions of future development.

Document visualization is a class of the information visualization techniques that transforms textual information such as words, sentences, documents, and their relationships into a visual form, enabling users to better understand textual documents and to lessen their mental workload when faced with a substantial quantity of available textual documents. [1]

And compared with text visualization that aims to visualize information on the text level, document visualization concentrates more on visualizing documents that include attributes and metadata except the core textual contents.

Document visualization has significant advantages over helping people to analyze and control big quantities of textual information in many cases. For example, we can intuitively get access to (1) word frequency or distribution; (2) semantic content and repetition; (3) the topic or topics that define document clusters; (4) the core content of document; (5) similarity among documents; (6) the connections among documents; (7) how content changes over time; and (8) information diffusion or other interesting patterns in social media, as well as improve text searches.

Generally ‘document’ is a textual record or physical form/representation of ‘information’.

The evolving notion of ‘document’ among Jonathan Priest, Otlet, Briet, Sch ¨ urmeyer, and the other documentalists increasingly emphasized whatever functioned as a document rather than traditional physical forms of documents. [2]

And with the development of digital technology, anything exists physically in a digital environment, such as a mail message or a technical report, could be considered as a document.

Documents are often minimally structured and may be rich with attributes and metadata, especially when concentrated in a specific application domain.

We may learn from the good practical guidelines to create an effective user interface for an interactive information visualization tool, as propounded by Ben Shneiderman who suggested in a form of mantra that an effective information visualization tool should follow the principle:
Overview first, zoom and filter, then details on demand. [4]

The mantra is accompanied by a task taxonomy for information visualizations that specifies seven
tasks at a high level of abstraction [4]:
• Overview. Gain an overview of the entire collection.
• Zoom. Zoom in on items of interest.
• Filter. Filter out uninteresting items.
• Details-on-demand. Select an item or group and get details when needed.
• Relate. View relationship among items.
• History. Keep a history of actions to support undo, replay, and progressive refinement.
• Extract. Allow extraction of sub-collections and of the query parameters.

We firstly divide document visualization methods into three main categories:
(1) single document visualization that has more emphasis on individual words and actual single document contents;
(2) document collection visualization that has more emphasis on large document collections, themes and concepts across collection, and how documents are relate to others;
(3) extended document visualization which often deals with comprehensive tasks, involves other attributes beyond the content of documents, and is always applied in specific field, such as social media and search.

In single document visualization, the goal is to quickly understand and absorb core content and text
features. The visualization focuses on words, phrases, semantic relations, and contents.

1. Vocabulary-Based Visualization
Vocabulary is the basic unit of a document. The visualization assists people in understanding words through visual representation of the document vocabulary features, such as word frequency, word distribution, and lexical structure, thereby providing a general idea of contents and features in a document.

Tag Clouds [6,7] and Wordle [8,9] are representative methods mainly visualizing word frequency. They are widely used in the news media and personal home pages. They provide layouts of raw tokens, colored, and sized by the corresponding word frequency within a single document. We may know the main research areas/content discussed in the text by the compact visual form of words.

Recently, some other methods have been proposed, extending the tag/word cloud, such as
parallel tag clouds (PTC), [10] ManiWordle, [11] context preserving dynamic word cloud,[12] visualization of internet discussion with extruded word clouds. [13]

Other examples: TextArc [14], DocuBurst [16].

2. Visualization Based on Semantic Structure

Visualization based on semantic structure usually use entities and their relationships to reveal the semantic content.

Semantic Graphs [19] is a visualization based on the semantic representation of a document in the form of a semantic graph. Firstly, it extracts subject–verb–object for each sentence by the Penn Treebank parse tree. Then, it links the triplets to their corresponding entity, which needs to resolve pronominal anaphors as well as to attach the associate WordNet synset. Thus, the document is summarized with the semantic graph and the list of extracted triplets.

3. Visualization Based on Document Content
Visualization based on document content is not only to search for specific words but also to obtain the characteristics and relations of the contents in the document.

The WordTree visualization provides the representation of both word frequency and context. Size is used to represent frequency of the term or phrase. The root of the tree is a user-selected word or phrase, and the branches represent the contexts in which the word or phrase is used in the document. Users can click on a branch, choose a different search term or re-center the tree. [23]

Martin Wattenberg’s Arc Diagrams [25] is a visualization method that focuses on showing complex patterns of repetition. It is suited to the analysis of highly structured data like musical compositions and less well-structured data like a web page. Repeated subsequences are identified and connected by semicircular arcs. Height of the arcs represents the distance between the subsequences; and thickness of the arcs represents the length of the subsequences.

Document Collection Visualization

Document collection visualization usually intends to reveal the topic or topics that define document clusters, the similarities and differences among documents, and how contents change over time.

1. Visualization of Document Themes

The main goal is to discover one or more specific topics and to reflect the relationships among various topics.

It may be used to find hot disciplines, evolutions, and trends.

The methods, such as ThemeScapes, [26] INSPIRE’s ThemeView, and The Galaxy, [27] all developed by the Pacific Northwest National Laboratory, having less emphasis on the time factor, focus more on characteristics of the document themes at some specific points.

ThemeView uses a 3D terrain map display to represent different themes. The height of a mountain represents the theme’s strength, and the distance between two mountains represents the similarity between the two themes. Keywords are used to distinguish each mountain. [27]

The Galaxy visualization uses a similar approach that themes are visualized as 2D clouds of document points-stars in a theme galaxy (Figure 8(b)). [27]

There are other representations for visualizing documents and topics as nodes in a node-link graph. TopicNets is a web-based system for visual and interactive analysis of large sets of documents using statistical topic models. [28] The main view is a document topic graph which can allow aggregate nodes. The time dimension is represented as a separate visualization, with documents placed chronologically around a broken circle, and connected to related topic nodes which are placed inside the circle.

The methods, such as ThemeRiver [29] and Topic Island, [30] have greater emphasis on the time factor, focusing more on visualizing thematic variations over time within a collection of documents.

ThemeRiver is in the form of axes, with the X-axis representing time and the Y-axis representing different themes. The ‘river’ flows from left to right through time, changing width to portray changes in theme strength of corresponding documents. Rivers of different colors represent different themes, and the width of river (i.e., narrow or wide) indicates the strength (decreasing or increasing) of an individual topic in the associated documents. [29]

2. Visualization of Document Core Content

Visualization of document core content mainly intends to give an overview of a collection of documents without reading them entirely.

Document Cards [35] visualizes large document collections, such as paper collections and news reports, which contain both texts and images to describe facts, methods, or stories. It represents the document’s key content as a mixture of images and important terms, similar to cards in a top trumps game. [35]

The pipeline for creating Document Cards is as follows: firstly, extract the text from the original document, and use a text mining approach to extract the key terms; then go to the phases of image extraction, including image processing and image packing; finally layout the extracted key terms and images to generate the corresponding document cards.

3. Visualization of Changes over Different Versions

Visualization of changes over different versions is used to visualize differences among multiple document versions that are generated over time.

History Flow [36] is designed to show changes between multiple document versions on Wikipedia. It can visualize the process of content changes and the corresponding authors who make the amendments. It also reveals some complex patterns of cooperation and confliction, such as vandalism and repair, anonymity versus named authorship, negotiation, and content stability.

Software visualization [37–39] focuses on visualizing the software development. SeeSoft, [40] Augur, [41] and Advizor [42] are visualizations for code documents. Xia gives visual insight into version control activities, like architectural and coding differences between two software versions. [43] Beagle visualizes changes among different released versions. [44] Spectrograph shows the time and location where changes happen in the system. [45]

4. Visualization of Document Relationships

With gradual increases in document quantity, the concepts and entities within documents become larger and larger, making the analyst’s task of evaluation and sense-making more difficult. Thus, it is quite meaningful to visualize connections among documents. The visualization focuses on the correlation among documents, like the connections among entities across different documents.

Jigsaw [46] is an interactive visualization for document exploration and sense-making, and it supports the analysis of relationships among documents. It visually shows connections between entities in the documents; where entities could be people, places, dates, organizations, and so on. It is suitable to documents describing a set of observations or facts, like news stories and case reports. It provides multiple views and each view provides a different perspective.

There are other methods for visualizing relations among multiple facets. ContexTour [47] presents the relations among conferences, authors, and topics in paper collections. FacetAtlas [48] shows relations among causes, symptoms, treatments, and diagnoses in Google Health documents. PivotPaths [49] visually explores relations of authors, keywords, and citations in academic publications.

5. Visualization of Document Similarity

In many cases of document collection visualizations, the goal is to place similar documents close to each other and dissimilar ones far apart.

The self-organizing map (SOM) [50] is a nonlinear projection method. It expresses complex, nonlinear relationships between high dimensional data items into simple geometric relationships on a 2D display.

When applying to information retrieval, it usually uses map displays. [51] Different colored areas represent different concepts in documents. Size of area indicates its relative importance in collection. Neighboring regions show commonalities in concepts. Dots in regions can represent documents. Additional information can be referred in Xia Lin’s map display [52] and WEBSOM. [53]

With the rise of social media (a textual medium), text streams, such as Twitter posts, are being generated in volumes that grow every day. A large body of research has appeared in recent years. Those works have different focuses and always involve multiple targets, such as dealing with the statistical analysis and presentation of topics or terms, focusing on the emergence of topic events, [33] and visualizing the text messages themselves. [54]

Christian Rohrdantz [55] provides an overview of real-time visualization of streaming text data.

STREAMIT [56] presents a similar visual representation of text streams which applies to news documents.

Whisper [57] fulfills the requirement for tracing information diffusion processes in social media, in a real-time manner.

Search Visualization visualizes the results of search operations. The relatively early approach is TileBars [58] that intends to minimize time and effort for deciding which documents to view in detail.

Susan Havre [59] introduces a graphical method called Sparkler for visually presenting and exploring the results of multiple queries simultaneously.

RankSpiral addresses the problem of how to enable users to visually explore and compare large sets of documents that have been retrieved by different search engines or queries. [60]

We have mainly considered the visualization objects and tasks when classifying document visualization methods. Our classification is considered more acceptable than other classifications (e.g., representations: pixel-based, map-based, tree-based graphs, node-link diagrams, circle graphs, etc.1), since visualization is usually task dependent, and users commonly begin with data and tasks. Actually, each method may belong to different category even under the same classification criteria; and we classify each method according to its key visualization focus (the visualization objects and tasks).

In Table 1, we summarize and compare those methods mainly from four aspects to give readers a brief view.

• Characteristics visualized. The characteristics of a document visualized by the method, as word frequency, semantic relations, content, changes, or connections among documents.

• Principles satisfied. The design principles satisfied, as noted in Document and Document Design section, the seven tasks: 1) Overview; 2) Zoom; 3) Filter; 4) Details-on-demand; 5) Relate; 6) History; 7) Extract.

• Requirements for a document. Document types suitable to the visualization method, i.e., whether the visualization method has special requirements for a document, like document content, structure, etc.

• Main features. Discuss the visualization method’s features, especially the versatility and interactivity.

Despite this, document visualization shares the same pipeline: get the data (a document or documents), transform it into vectors, then run algorithms based on the tasks of interest (i.e., similarity, search, clustering) and generate the visualizations.


Document visualization techniques combine human wisdom and computer graphics, allowing users to efficiently and intuitively browse, explore, and understand the increasing quantity of documents.

1. Extension: Existing methods can be extended to suit for large-scale document collections.
2. Versatility: It is significant to design relatively general visualization models for different tasks within this field, since existing methods always have narrow scope of application due to its pointed direction.
3. Interactivity: It is important to design a more intuitive man–machine interface to improve user’s experience of interaction. Also it is crucial to find some interstices to allow users to participate in researching and developing process, especially the testing period.
4. Techniques:
• Algorithms. Develop and adopt scalable, high-performance algorithms for text processing, such as text summary and clustering.
• Parallel processing technology. With the adoption and popularity of Coordinated and Multiple Views (CMV), a visualization system usually includes multi-views.
• Real-time processing technology.

1. Evaluation Many document visualization methods or even information visualization methods lack a quantitative measurement which can indicate the overall quality, novelty, uncertainly, and other evaluative metrics. More recently, there exist more and more publications that reflect upon current practices in visualization evaluation. In fact, the BELIV workshop was created as a venue for researchers and practitioners to ‘explore novel evaluation methods, and to structure the knowledge on evaluation in information visualization around a schema’. [62]

2. Theoretical Foundations The 2007 Dagstuhl Workshop identified collaborative information visualization with theory building as major directions for future development. [63]