顯示具有 influential authors 標籤的文章。 顯示所有文章
顯示具有 influential authors 標籤的文章。 顯示所有文章

2015年12月18日 星期五

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

近年來由於可以取得大量而多樣的文本資料和採用文本處理演算法等原因,研究人員對文本視覺化(text visualization)與視覺性的文本解析(visual text analytics)的研究興趣增加。本研究針對文本視覺化技術提出一個互動的視覺調查(visual survey)。並且利用此次調查的資料,分析文本視覺化的現況,比較研究使用的各種分析與視覺化技術,以及分析有關研究者的資訊,以提供搜尋相關研究、探索次領域(subfield)以及獲得研究趨勢的洞察等目的

本研究採納前人的研究,將文本視覺化技術,以分析任務(analytic tasks)、視覺化任務(visualization tasks)、資料領域(data domain)以及資料來源(data source)、資料性質(data property)、視覺化的維度(visualization dimensionality)、視覺化的呈現(visualization representation)、視覺化的排列方式(visualization alignment)等面向,建立分類架構(taxonomy)。


分析任務是指使用者採用文本視覺化技術預期達到的主要目的,這些分類包括:
1. 文本摘要 (Text Summarization) / 主題分析 (Topic Analysis) / 實體抽取 (Entity Extraction)
2. 言談分析 (Discourse Analysis):文本或對話轉錄(conversation transcript)裡流動的語言學分析。
3. 情感分析 (Sentiment Analysis)
4. 事件分析 (Event Analysis)
5. 趨勢分析 (Trend Analysis) / 樣式分析 (Pattern Analysis)
6. 詞法/語法分析 (Lexical / Syntactical Analysis)
7. 關係/連結分析 (Relation / Connection Analysis)
8. 翻譯/文本比對分析 (Translation / Text Alignment Analysis)

視覺化任務則是由文本視覺化技術所支援的較基層呈現與互動任務,包括:

1. 自動凸顯/建議興趣區 (Region of Interest)
2. 群集 (Clustering) / 分類 (Classification / Categorization)
3. 比較 (Comparison)
4. 概觀 (Overview)
5. 監視 (Monitoring)
6. 瀏覽 (Navigation) / 探索 (Exploration)
7. 對於不確定的對策 (Uncertainty Tackling)

資料領域,包括

1. 線上社交媒體 (Online Social media)
2. 通訊 (Communication)
3. 專利 (Patents)
4. 評論 (Reviews) / 病歷 (Medical Records)
5. 文學作品 (Literature) / 詩 (Poems)
6. 科學文章 (Scientific Articles) / 論文 (Papers)
7. 社論媒體 (Editorial Media)

資料來源有單一文件 (Document) [33]、語料庫 (Corpora) [25]以及 串流文本 (Streams) [19];特殊的資料性質包括地理空間 (Geospatial) [11]、時間序列 (Timeseries) [14] 以及網路 (Networks) [6];視覺化的再現包括下列項目:折線圖 (Line Plot) / 河流圖 (River) [9, 18]、像素 (Pixel) / 面積 (Area) / 矩陣 (Matrix) [13, 7, 4]、節點-連結 (Node-Link) [32]、雲 (Clouds) / 銀河 (Galaxies) [1, 3]、地圖 (Maps) [34]、文本 (Text) [26]與形符 (Glyph) / 圖標 (Icon) [28, 10];排列則包括了輻射狀 (Radial) [35]、線性 (Linear) / 平行線 (Parallel) [8] 以及測標依賴 (Metric-dependent) [22]。

本研究指出有超過一半(56%)的文本視覺化利用主題模型(topic modeling)技術,資料來源方面大多數支援語料庫(70%),並且許多支援時間相關的資料(43%),而視覺再現方面主題以二維(2-D)為主,僅有極少數的研究以三維(3-D)的方式呈現,約占所有研究的4%。

文本視覺化的前五位主要作者為Daniel A. Keim (17 筆)、Shixia Liu (12 筆)、Christian Rohrdantz (9 筆)、Daniela Oelke (7 筆)和 Huamin Qu (7 筆)。將作者依據他們的合著關係建立研究者合作網路圖後,觀察網路圖的相連成分,可以發現大部分是獨立的小群體,最大的成分上共有106位作者,並且在這個成分上的兩個主要集群為University of Konstanz和Microsoft Research Asia等兩個研究團隊,Daniel A. Keim 和 Shixia Liu分別為集群的中心,並且他們二位也是網路圖上中介中心性最高的節點。雖然在本研究蒐集的資料上,這兩位作者之間並沒有直接的合作關係,但他們都曾與中介中心性第三高的兩位作者Dongning Luo 和 Jing Yang合作。


In this paper, we present an interactive visual survey of text visualization techniques that can be used for the purposes of search for related work, introduction to the subfield and gaining insight into research trends.

The interest for text visualization and visual text analytics has been increasing for the last ten years. The reasons for this development are manifold, but for sure the availability of large amounts of heterogeneous text data (caused by the popularity of online social media) and the adoption of text processing algorithms (e.g., for topic modeling) by the InfoVis and Visual Analytics communities are two possible explanations.




Analytic Tasks
these items are critical to the main analysis goals that users expect to achieve when employing a text visualization technique.

1. Text Summarization / Topic Analysis / Entity Extraction

2. Discourse Analysis
the linguistic analysis of the flow of text or conversation transcript.

3. Sentiment Analysis
for techniques related to the analysis of sentiment, opinion, and affection.

4. Event Analysis
deal with the extraction of events from the text data or involve visualization of text in some different manner

5. Trend Analysis / Pattern Analysis
both automated trend analysis and manual investigation directed at discovering patterns in the textual data.

6. Lexical / Syntactical Analysis

7. Relation / Connection Analysis

8. Translation / Text Alignment Analysis

Visualization Tasks
lower-level representation and interaction tasks that are supported by the text visualization techniques.

1. Region of Interest
the automatic highlighting/suggestion of data items/regions that could be of interest to the user for more detailed investigation

2. Clustering / Classification / Categorization

3. Comparison

4. Overview
both techniques that provide “the big picture” by displaying a significant portion of the data set as well as techniques which use special aggregated representations to provide overview while reducing the visual complexity

5. Monitoring

6. Navigation / Exploration

7. Uncertainty Tackling


Domain

1. Online Social media

2. Communication

3. Patents

4. Reviews / (Medical) Records

5. Literature / Poems

6. Scientific Articles / Papers

7. Editorial Media

Data sources include the following self-evident items: Document [33], Corpora [25], and Streams [19].

The special data properties include Geospatial [11], Timeseries [14], and Networks [6].

Representation includes the following items: Line Plot / River [9, 18], Pixel / Area / Matrix [13, 7, 4], Node-Link [32], Clouds / Galaxies [1, 3], Maps [34], Text [26], and Glyph / Icon [28, 10].

Alignment, i.e., layout, includes Radial [35], Linear / Parallel [8], and Metric-dependent [22].

As displayed in the table, our proposed taxonomy includes most of the categories except for two: we believe that the underlying data representation (e.g., bag-of-words vs. language model [30] or whole text vs. partial text [24]) is more relevant to the underlying computational methods than to observable visualization techniques.

And the same naturally holds for data processing methods (e.g., the specification of involved MDS methods [2]) that are partially covered by other categories in our taxonomy, for instance, the analytic task of topic analysis implies the usage of corresponding computational methods.

Using the data collected for the survey, we have been able to analyze the general state of the text visualization field, to compare the usage of various analysis and visualization techniques (with regard to our taxonomy), and to analyze the information about researchers in this field.

According to our current set of entries, the trend for rapid increase of text visualization techniques started around 2007.

With regard to category statistics (cf. Fig. 4), there is an obvious interest for tasks related to topic modeling (56% of all entries).

The majority of the techniques support corpora as data sources (70% of all entries), and a lot of them support time-dependent data (43% of all entries).

Another result—which is probably expected—is that only less than 4% of all entries use 3-dimensional visual representations.

We have also taken a look at the authorship statistics for the current data set. The top five authors with regard to number of techniques are Daniel A. Keim (17 entries), Shixia Liu (12 entries), Christian Rohrdantz (9 entries), Daniela Oelke (7 entries), and Huamin Qu (7 entries).

As seen in Fig. 5, the majority of author nodes are included into isolated connected components of small sizes (less than 10 nodes) while there is a big connected component with 106 nodes present in the graph.

The two major clusters in that component represent the research groups from the University of Konstanz and Microsoft Research Asia with Daniel A. Keim and Shixia Liu as cluster center nodes.

Shixia Liu and Daniel A. Keim happen to have the 1st and the 2nd largest betweenness values in the graph, respectively. While these two researchers have no direct collaboration with regard to our data set, they both have collaborated with Dongning Luo and Jing Yang who both share the 3rd largest betweenness value.

2014年1月18日 星期六

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

本研究利用社會網絡分析、共現分析(co-occurrence analysis)、叢集分析和詞語的頻率分析等多種分析技術,從Scientometrics期刊1978到2004年發表的1927筆論文資料,探討科學家合作網絡的結構特性、整個網絡上的合作領域以及個別的合作網絡、合作網絡上的合作中心(collaborative  center)。

過去的研究裡,Schubert (2002) 和 Dutt, Garg, & Bali (2003)都是針對國家間合作的巨觀層次。Kretschmer (2004) 認為巨觀和中觀(meso)層次的分析無法足夠地反映個人之間的合作趨勢,因此呼籲應在微觀層次的分析投注更多努力。

1927筆論文資料裡,單一作者的論文共有1052筆,所以仍稍占多數。作者數大於3的論文僅占非單一作者論文的13.71% (120/875),顯然研究Scientometrics的團隊規模都不大。發表3篇論文以及以上的高生產作者共計234人,其中有69.66%的作者曾發表與其他作者合作的論文。將這些作者間的合作關係表現成網絡,並利用Bibexcel對這個網絡上的節點進行叢集分析,共發現22個叢集。前兩個較大的叢集分別有15與14個科學家。網絡上最大的相連成分上共有15個叢集,共有合作經驗的高生產作者中的96位,占58.90%。合作網絡共有401條連結線,網絡密度為0.03,顯示Scientometrics領域的合作很鬆散。

對每一個節點計算它們的三種中心性,結果發現中心性和對應作者的生產力之間有很顯著的正相關,表示高生產力的作者同時也活躍在Scientometrics領域的合作網絡上。其中Glänzel的程度中心性最高,總共和其他18位作者有合作關係。

以詞語的頻率分析每個叢集的主題,最大的兩個叢集有類似的主題,但使用的研究方法略有不同。此外,研究主題為科學合作的四個叢集間幾乎沒有連結,同樣的情形也發生在研究科學與技術之間關係的四個叢集。

The structure of scientific collaboration networks in scientometrics is investigated at the level of individuals by using bibliographic data of all papers published in the international journal Scientometrics retrieved from the Science Citation Index (SCI) of the years 1978–2004.

Combined analysis of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of words is explored to reveal: (1) The microstructure of the collaboration network on scientists’ aspects of scientometrics; (2) The major collaborative fields of the whole network and of different collaborative sub-networks; (3) The collaborative center of the collaboration network in scientometrics.

Schubert [8] and Dutt etc. [9] presented international collaboration characteristics in the scientometrics community itself, focusing on country aspects at macro level.

Kretschmer [6] appealed to devote more efforts to investigations at micro level in the future because the knowledge at meso and macro level does not yet adequately reflect the trends in cooperation between individuals.

The study is based on bibliographic data retrieved from the Web of Science. The data contains all types of documents published in Scientometrics during 1978 to 2004.

In this study we have adapted an integrated procedure of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of title words.

Bibexcel is designed as a tool for manipulating bibliographic data, which is a free online-software published by Persson. In the present study, Bibexcel is used to do cooccurrence analysis and cluster analysis.

Following the methods of Otte & Rousseau [11], White [13] and Kretschmer & Aguillo [12], SNA was applied to display the microstructure of collaboration networks in scientometrics with Pajek.

Moreover, we used frequency analysis of title words to display the main collaborative field of different sub-networks. The software for frequency analysis is demo version of Wordsmith Tools published by Oxford University Press and available online.

There were 1927 documents published in Scientometrics during 1978 to 2004 (see Table 1).



From Table 1, we found that the pattern of co-authorship was still dominated by single-authored papers as the conclusion drawn by Dutt etc. [9].

While the number of multi-authored papers (the number of co-authors is more than 3) accounts for 13.71% only, which indicates that team size in scientometrics is not large.

In order to show the main structure of the network, each author must published 3 papers or more to be included in this integrated analysis. This threshold resulted in a total of 234 prolific authors publishing 3 or more papers during 1978 to 2004, among them there are 163 authors published co-authorship papers, accounting for 69.66% of the prolific authors.



Based on cluster analysis embedded in Bibexcel, we gained 22 clusters circled by solid lines (see Figure 1). We identified these clusters as sub-networks in the field of scientometrics.

The largest subnetwork is number 1 that has 15 collaborators, and the second largest one is number 2, which has 14 collaborators, and so on.

We noticed that there was totally 15 subnetworks connected with each other composing the largest central component, which had 96 numbers accounting for 58.90% of the prolific authors published co-authorship papers.

Density is an indicator for the general level of connectedness of the graph. ... In the present study, there are totally 401 links in the network, so the density of the network is 0.03, which indicates that the collaborative network in the field of scientometrics is very loose.

So an author who has high degree centrality must has collaborated with many other authors, which means the author is a central collaborator of the whole network. In the present study, Glänzel who has 18 co-workers is the central author of the whole network.

We found a positive and significant correlation between output of authors and the centrality measures (r=0.648, 0.437, 0.338 respectively at the 0.01 level, see Table 4) after investigating the correlations between output and the three centralities of the 125 authors in the 22 sub-networks, which indicated that most of the prolific authors are also active in collaboration network in the field of scientometrics.

We have also presented the main collaborative field of different sub-networks in scientometrics and found that the two biggest sub-networks have the similar collaborative topic with slightly methodological difference. In addition, we found an interesting phenomenon that four sub-networks dealing with scientific collaboration didn't collaborate with each other except sub-network 3 and 12. Moreover, four subnetworks studying technology and science never collaborated with each other at all.

2014年1月15日 星期三

Chen, Y., Börner, K., & Fang, S. (2013). Evolving collaboration networks in Scientometrics in 1978–2010: a micro–macro analysis. Scientometrics, 1-20.

Chen, Y., Börner, K., & Fang, S. (2013). Evolving collaboration networks in Scientometrics in 1978–2010: a micro–macro analysis. Scientometrics, 1-20.

科學計量學(Scientometrics)利用數學、統計與資料分析方法與技術,蒐集、處理、解釋與預測學術傳播、成效、發展與動態等科技的特徵,對科技進行量化研究。就實務的技術而言,科學計量學利用書目計量的概念測量文本與資訊,並且利用科學地圖(science map)展現結果 (Börner 2010; Börner et al. 2003)。本研究利用網絡分析技術,從巨觀(國家)、中觀(機構)與微觀(作者)三種層次,探討Scientometrics期刊1978-2010年發表的2541筆論文上的合作情形。

過去對於Scientometrics有以下的相關研究,Schoepflin and Glänzel (2001) 利用1980、1989和1997三年出版的Scientometrics論文,發現科學政策(science policy)與科學社會學(the sociology of science)等主題相關的論文比率減少。Peritz and Bar-Ilan (2002) 以1990和2000年的Scientometrics論文,確認Research Policy和Social Studies of Science分別是第三和第四最常引用的期刊。Hou, Kretschmer, and Liu (2008)從Scientometrics2002到2004年論文上的作者合作網絡上發現一半以上的作者有合作的經驗,但是網絡的連結並不強而且疏鬆。Dutt, Garg, and Bali (2003)使用大量論文資料進行,研究資料期間為1978到2001,發現機構的平均論文數偏低,顯示研究的產出相當分散,而且以單一作者的論文為主,雖然多位作者的論文正蓄勢待發。

研究結果發現:
(1) 論文生產力較大的國家有美國、比利時、英國、荷蘭和西班牙。
(2) 機構與作者數隨時間增加,但是機構的平均論文數成長緩慢,近年的作者平均論文數則減少。
(3) 具有高中心性及中介性的一些機構可視為是合作網絡上的守門人(gatekeepers)。
(4) 近期較高生產力的作者取代了早期的重要作者。

Scientometrics的論文、作者和機構平均每年增加率為20%,顯示這個領域吸引愈來愈多的研究人員和機構加入。另外,Bettencourt et al. (2009) 以下面的公式指出當領域成長時,它的合作網路將會變得更為稠密。

而本研究三種層次的合作網絡,國家合作網絡的α值為 2.9533,機構與作者則分別為1.5222和1.2353。明顯的可以看出國家合作網絡相當快速地變得稠密,但由於單一作者論文的增加以及許多合作仍然是同一國家內的機構或同一機構內的作者彼此間的合作,使得機構合作網絡和作者合作網絡的α值較小。
在網絡的直徑(diameter),也就是網絡上最長的路徑方面,國家合作網絡的直徑在1989到1998年是增長的情形,但在近十年則是減短;反之,機構合作網絡和作者合作網絡到2010年仍在持續增長。
測量三種層次合作網絡的節點的連結程度(connection degree),其分布情形都符合冪次法則(power law)。節點的重要性可以從它們的程度中心性和中介中心性來推測,早期和近期的重要作者有很大的不同,在國家與機構方面的差別相當小。
比較三種層次的結果可以發現,較低層的結果會影響到上面的層次。例如有些作者的高排名不僅影響機構的排名,同時也會主導國家的排名;當具有高生產力的作者移動後,會牽動機構網絡的結構性變化。

Specifically, we would like to understand if and how collaborations at the author (micro) level impact collaboration patterns among institutions (meso) and countries (macro).

All 2,541 papers (articles, proceedings papers, and reviews) published in the international journal Scientometrics from 1978–2010 are analyzed and visualized across the different levels and the evolving collaboration networks are animated over time.

(1) USA, Belgium, and England dominated the publications in Scientometrics throughout the 33-year period, while the Netherlands and Spain were the subdominant countries;

(2) the number of institutions and authors increased over time, yet the average number of papers per institution grew slowly and the average number of papers per author decreased in recent years;

(3) a few key institutions, including Univ Sussex, KHBO, Katholieke Univ Leuven, Hungarian Acad Sci, and Leiden Univ, have a high centrality and betweenness, acting as gatekeepers in the collaboration network;

(4) early key authors (Lancaster FW, Braun T, Courtial JP, Narin F, or VanRaan AFJ) have been replaced by current prolific authors (such as Rousseau R or Moed HF).

Comparing results across the three levels reveals that results from one level might propagate to the next level, e.g., top rankings of a few key single authors can not only have a major impact on the ranking of their institution but also lead to a dominance of their country at the country level; movement of prolific authors among institutions can lead to major structural changes in the institution networks.

Scientometrics is a distinct discipline that performs quantitative studies of science and technology using mathematical, statistical, and data-analytical methods and techniques for gathering, handling, interpreting, and predicting a variety of features of the science and technology enterprise, including scholarly communication, performance, development, and dynamics.

In practice, scientometrics often requires the use of bibliometrics, the measurement of texts and information, and results might be presented as science maps (Börner 2010; Börner et al. 2003).

The study presented here uses papers that appeared in Scientometrics, the flagship journal of the field (Chen et al. 2002) publishing a major percentage of works in scientometrics as well as in the field of informetrics (Bar-Ilan 2008) over the last 33 years.

For example, Schoepflin and Glänzel (2001) used papers published in Scientometrics for the years 1980, 1989, and 1997 to identify a decrease in the percentages of both the articles related to the subjects of science policy and to the sociology of science.

Peritz and Bar-Ilan (2002) used papers published in Scientometrics for the years 1990 and 2000 and confirmed that Research Policy and Social Studies of Science are the third and fourth most frequently referenced journals in articles published in Scientometrics.

Hou et al. (2008) analyzed the structure of scientific collaboration networks in scientometrics at the micro level (individuals) by using bibliographic data of all papers published in Scientometrics from the years 2002–2004. They found that although half the authors had co-authored with each other, the network was not strongly connected and the collaborative network in the field of scientometrics was very loose.

Dutt et al. (2003) analyzed Scientometrics papers published during 1978–2001, examining the distribution of countries and themes and comparing institutions and coauthors to show that the research output is highly scattered, as indicated by the average number of papers per institution and dominated by single-authored papers; however, multi-authored papers are gaining momentum.

Chen et al. (2010) introduced a multiple-perspective co-citation analysis for characterizing and interpreting the structure and dynamics of co-citation clusters of the field of information science between 1996 and 2008. He showed that the multiple-perspective method increases the interpretability and accountability of both author-citation analysis (ACA) and document- citation analysis (DCA) networks.

Wagner and Leydesdorff (2005) applied network analysis to map the growth of international co-authorships, and they found that international co-authorships can be explained based on the organizing principle of preferential attachment, although the attachment mechanism deviates from an ideal power-law.

Samoylenko et al. (2006) visualized the scientific world and its evolution by constructing minimum spanning trees (MSTs) and a two-dimensional map of scientific journals using the Science Citation Index from the Web of Science database for 1994–2001 and showed a linear structure of the scientific world with three major domains: physical sciences, life sciences, and medical sciences.

Perc (2010) studied the evolution of Slovenia’s scientist collaboration network from 1960 to 2010 with a yearly resolution and showed the network had a ‘‘small world’’ pattern and its growth was governed by near-linear preferential attachment. This paper will advance the existing works by studying the evolution of scientometrics at three different network levels.

Figure 1 shows the growth (annual and cumulative) of the number of papers, countries (or regions), institutions and authors from 1978 to 2010. By counting the annual numbers in each figure, we obtain average annual growth rates, which are 20.4 % (papers), 9.4 % (countries), 19.6 % (institutions), and 20.1 % (authors).

As Bettencourt et al. (2009) pointed out, when fields grow, their collaboration networks densify—i.e., the average number of edges per node increases over time. They found that the relation between the number of nodes and edges followed a simple scaling law with scaling exponent (α > 1):



Figure 2 shows that the scaling exponent a equals 2.9533 at the macro-country, 1.5222 at the meso-institution, and 1.2353 at the micro-author levels. It has the highest value for countries—i.e., the country collaboration networks densify rather quickly, which is also due to the fact that this is the network with the fewest nodes. However, a large number of within-country or within-institution collaborations or an increase in single-authored papers would also result in smaller α values.

The diameter of a collaboration network has major implications for information diffusion—the shorter a pathway of coauthor linkages that connects an author pair, the more likely knowledge diffuses.

Over the 33 years, the country collaboration network diameter grew from 1989 to 1998 (there were no edges before 1989), achieves the highest value in 1998, and decreases in the last 10 years. This might be due to the rather limited number of countries that perform scientometrics research.

The diameters of the institution and author collaboration networks increase continually and both reach a diameter d = 15 in 2010.

A closer look at the density of the three networks (the ratio of the number of actual edges to all possible edges in a fully connected graph with the same number of nodes) shows that both the meso and micro networks’ densities decrease over time while the macro network, which experienced a topological transition from large to decreasing diameter, shows an increase in density.

In an attempt to understand the structure of the 1978–2010 networks, the degree for each node in the network was determined and the node degree distribution p(k) plotted in Fig. 3. ... All three networks exhibit power law degree distributions.

To understand which countries, institutions, and authors play key roles in the three networks, the degree centrality (the number of links a node has) and betweenness centrality (nodes that have a high probability to occur on a randomly chosen shortest path between two randomly chosen nodes have a high betweenness) (Freeman 1977) values for each node were calculated. The resulting TOP-5 countries, TOP-10 institutions, and TOP-10 authors calculated for every 6 years (cumulatively from 1978) are listed in Tables 1 and 2.

In addition, the last table column shows the TOP-10 countries, institutions, and authors if only 2001–2010 data is considered. While the differences are minimal for countries and institutions, the list of TOP-10 authors changes considerably if only recent works are considered.

Figure 4 shows that, by the end of 2010, Belgium, USA, England, Germany, the Netherlands, China, and France are central network nodes with a large number of papers. These six countries not only link to each other but also to outside countries—e.g., Belgium and Germany have strong links to Hungary, and Belgium and England have strong links to Finland.

When analyzing the evolving institution collaboration networks, it becomes clear that a few key institutions manage to stay in the TOP-10 list—among them are the Univ Sussex, KHBO, Katholieke Univ Leuven, Hungarian Acad Sci, and Leiden Univ.

During the evolution of the co-author networks, early authors are replaced by current authors. Most TOP-10 authors from 1980 and 1986 are missing in the later years. Key authors listed in the TOP-10 lists around 1986 decline in ranking or are replaced by other authors.

One might assume that rankings on the author (micro) level impact the ranking of institution (meso) and country (macro) levels. While author rankings impact institution rankings; institution rankings are less predictive of country rankings, as exemplified below.

As can be seen in Table 3, USA ranks first in the number of institutions and the number of papers over the 33 year time span. However, the average number of papers per institution was low for the USA, especially when compared with Belgium, Netherlands, and Hungary. ... Similarly, while no single author in the USA appears in the TOP-10 lists, the number of all authors combined and the number of their papers results in a high country ranking.

Can one single author impact the ranking of an entire institution or country? The answer is yes. ... The 155 papers of the Hungarian Academy of Sciences were co-authored with 30 institutions, 22 of which were contributed by papers authored by Glänzel W. As for the 93 papers by the Katholieke Universiteit Leuven, 13 of 51 institution links were added by Glänzel W.

Over the 33 years, the number of countries grew steadily with a linear growth feature with USA, Belgium and England leading in terms of centrality and betweenness. ... As their share increases, they have a stronger impact on the evolution of scientometrics. Over time, more and more collaboration links are generated and the average node degree and network density increase as well (see Table 4).

It is important to point out that some top-ranking countries have a small number of top-ranking institutions (e.g., Katholieke Univ Leuven in Belgium) while other countries (USA) have a large number of contributing institutions.

Similarity, some top-ranking institutions have one or two top-ranking authors, e.g., Glänzel W and Rousseau R

That is, single authors can not only have a major impact on the ranking of their institution but also of their country.

At the same time, the growth rate of institutions, authors and papers for each year were similar about 20 %. It suggested that this field had been attracting more and more institutions and authors to join the field of scientometrics.

The co-author network analysis showed that many new authors joined the field of scientometrics, especially in the recent 8 years. The diameter, average degree, and density of the network show the same trends as those calculated for institutions.

While co-author networks experience the departure of senior and the arrival of young researchers, the institution and country networks seem to have a comparatively stable structure of key nodes.

2013年11月29日 星期五

Zhang, H., Qiu, B., Giles, C. L., Foley, H. C., & Yen, J. (2007, May). An LDA-based community structure discovery approach for large-scale social networks. In Intelligence and Security Informatics, 2007 IEEE (pp. 200-207). IEEE.

Zhang, H., Qiu, B., Giles, C. L., Foley, H. C., & Yen, J. (2007, May). An LDA-based community structure discovery approach for large-scale social networks. In Intelligence and Security Informatics, 2007 IEEE (pp. 200-207). IEEE.

本研究以LDA (latent Dirichlet allocation) 演算法[17]做為社群發現(community discovery)的方法,也就是找出社會網路上在群體內高度相連但在群體間相對稀疏的行動者(actors)社群。在這個方法裡,社群被視為是LDA模型裡的隱藏變數,由所有行動者依不同比例組合成的混合(mixtures)。本研究認為這個方法的優點是只需要利用網路的型態資訊(topological information)便可進行社群發現;此外,有別於[2, 3, 5]等先前的方法,這個方法能夠描述行動者屬於多個社群的現象,並且在每個社群上有不同重要性。

本研究將所有與某個行動者互動的行動者以及他們之間的社會互動情形定義為這個行動者的社會互動特徵(social interaction profile, sip),例如行動者vi的社會互動特徵定義為

SIW(viwij)是行動者vi與另一位行動者wij間的互動情形,mi是與vi互動的行動者數目。
在利用LDA模型進行社群發現時,將每個行動者的社會互動特徵視為是原先模型中的文件,這些行動者發生的互動則是模型中的詞語,值得一提的是在社會網路上行動者間的互動順序是可以交換的(exchangeable),因此LDA模型相當適合這種情形。下圖是這個模型的圖示,


對行動者vi而言,其社會互動特徵sipi的產生過程(generative process)如下
1) Sample mixture components ϕk ~ Dir(β) for k (belongs) [1, K]
2) Choose θi ~ Dir(α)
3) Choose Ni ~ Poisson(ξ) (note that Poisson assumption is not critical to this model)
4) For each of the Ni social interactions wij :
(a) Choose a community ιij ~ Multinomial(θi);
(b) Choose a social interaction wij ~ Multinomial(ϕιij)
也就是
1) 從Dir(β)中產生行動者的組合ϕk來代表K個社群中每一個社群;
2) 從Dir(α)中選擇一個社群的組合θi,做為所有與vi互動的行動者所可能屬於的社群的分布情形;
3) 從Poisson(ξ)中選擇社會互動的個數Ni ;
4) 產生每一個社會互動 wij時:
a) 從Multinomial(θi)中選擇一個社群ιij
b) 從Multinomial(ϕιij)中選擇一個社會互動 wij

根據這個模型,行動者vi的社會互動特徵sipi中的第j個社會互動元素wijwm的機率是

此處θisipi 的混合比變數(mixing proportion variable),而ϕk是第k個社群成分分布(component distribution)的參數集合。

給定超參數(hyperparameter)αβ後,所有已知和隱藏變數的聯合機率(joint probability)為

由於要精確地推論LDA模型的變數一般而言是很困難的(intractable),因此經常以 變異期望值最大化(variational expectation maximization) [17], 期望值延遲 (expectation propagation) [24]和 Gibbs取樣 (Gibbs sampling) [25], [26], [22]等三種方法取得近似解。本研究採用的方法是Gibbs取樣,以Markov鏈 Monte Carlo模擬(Markov-chain Monte Carlo simulation)的方式,去逼近ϕk,wθm,k等變數


本研究針對CiteSeer和NanoSCI兩個書目資料集中的作者合著網路(co-authorship networks)上最大的成分進行社群發現,其中CiteSeer部分共包含249866個節點,而NanoSCI則有203762個節點。本研究並且提出了01-SIP、012-SIP和 k-SIP等三種描述行動者間的社會互動情形,01-SIP是描述兩個作者有直接合著的社會互動情形,如果兩個作者曾經合著至少一篇論文,便將他們的互動情形設為1,否則便設為0;012-SIP則考慮兩個作者沒有直接合著但曾經分別和第三位作者合著的情形,如果兩個作者曾經合著至少一篇論文,便將他們的互動情形設為2,雖然這兩位作者不曾合著,但曾有分別和第三位作者合著,便將互動情形設為2,否則便設為0;k-SIP則描述兩位作者多次合著的情形,k是他們合著的次數。研究結果發現在複雜度(perplexity)與各產生社群的緊密度(compactness of communities)等指標,012-SIP的成效都比其他兩者好,從社群發現的結果可以發現許多社群裡的作者是屬於相同的研究機構或是彼此具有相似的研究興趣,因此會共同研究並且合作發表而成為社群。

This paper describes an LDA(latent Dirichlet Allocation)-based hierarchical Bayesian algorithm, namely SSN-LDA(Simple Social Network LDA). In SSN-LDA, communities are modeled as latent variables in the graphical model and defined as distributions over the social actor space. The advantage of SSN-LDA is that it only requires topological information as input.

This model is evaluated on two research collaborative networks: CiteSeer and NanoSCI. The experimental results demonstrate that this approach is promising for discovering community structures in large-scale networks.

An important task in these emerging networks is community discovery, which is to identify subsets of networks such that connections within each subset are dense and connections among different subsets are relatively sparse.

Unlike those previous community discovery studies, we design a hierarchical Bayesian network based approach, namely SSN-LDA (Simple Social Network-LDA) to discover probabilistic communities from social networks. ... In this model, communities are modeled as latent variables and are considered as distributions on the entire social actor space.

We also propose three different approaches to create social interaction profiles based on the social interaction information in the network.

Girvan et al extended this measure to edges and designed a clustering algorithm which gradually remove the edges with highest betweenness value [5]. ... However, a major problem with this approach is that the complexity of this approach is O(m2n), where m is the number of edges in the graph and n is the number of vertices in the network.

The graph partition problem can be formulated as the balanced minimum cut problem where the goal is to find an optimal graph partition so that the edge weight between the partitions is minimized while maintaining partitions of a minimal size. The NP-complete complexity of this approach [15] requires approximate solutions. Flake et al developed approximate algorithms to partition the network by solving s-t maximum flow techniques [2], [3]. The main idea behind maximum flow is to create clusters that have small inter-cluster cuts and relatively large intra-cluster cuts.

The major difference between SSN-LDA approach and the aforementioned approaches is that SSN-LDA is a mixture-model based probabilistic approach. Each community weighs in the contributions from every social actors and this property can be exploited in many potential applications that will be introduced in Section VI.

LDA model was first introduced by Blei for modeling the generative process of a document corpus [17].

Among these variants of LDA models, the approaches proposed in [20], [10] are both concerned about the authors of the documents in the corpus. In particular, Zhou et al introduced a community latent variable in their graphical model and applied it to discover community information embedded in document corpus. This approach can discover the underlying social network based on social interactions and topical similarity.

In their follow-up work[23], Zhou et al. investigated how research topics evolve over time and attempted to discover the most influential researchers involved in such transitions.

Each actor is characterized by its social interaction profile (SIP), which is defined as a set of neighbor(wij) and the corresponding weight(SIW(vi, wij)) pair.

where mi is the size of vi's social interaction profile.

Note that we consider the social interaction elements in this profile are exchangeable and therefore their order will not be concerned. It is this exchangeability that permits the application of LDA model [17].

Subsequently, we specify that a social network contains a set of communities ι (ι1, ι2, …, ιk) and each community in ι is defined as a distribution on the social actor space. In SSN-LDA, community assignments are modeled as a latent variable (ι) in the graphical model. The community proportion variable (θ) is regulated by a Dirichlet distribution with a known parameter α. Meanwhile, each social actor belongs to every community with different probabilities and therefore its social interaction profiles can be represented as random mixtures over latent communities variables.

The SSN-LDA model for social network analysis is illustrated in Fig. 1. Note that SSN-LDA resembles topic-based LDA model[17], with the social network being analogous to the corpus, the social interaction profiles being analogous to documents; and the occurrence of social interactions being analogous to words.

The distribution of topics in documents and the terms over topics are two multinomial distributions with two Dirichlet priors, whose hyperparameters are α and β respectively.

The dimensionality K of the Dirichlet distribution, which is also the number of community component distributions, is assumed to be known and fixed.

This generative process for an agent(wi)'s social interaction profile sipi in a social network is:
1) Sample mixture components ϕk ~ Dir(β) for k (belongs) [1, K]
2) Choose θi ~ Dir(α)
3) Choose Ni ~ Poisson(ξ) (note that Poisson assumption is not critical to this model)
4) For each of the Ni social interactions wij :
(a) Choose a community ιij ~ Multinomial(θi);
(b) Choose a social interaction wij ~ Multinomial(ϕιij)

According to the model, the probability that the jth social interaction element wij in the social actor wi's social interaction profile sipi instantiates a particular neighboring agent wm is:

where θi is the mixing proportion variable for sipi and ϕk is the parameter set for the kth community component distribution.

Given the hyperparameters α and β, the joint distribution of all known and hidden variables is:


Exact inference is generally intractable for LDA model. There have been three major approaches for solving this model approximately, including variational expectation maximization [17], expectation propagation [24], and Gibbs sampling[25], [26], [22].

Gibbs sampling is a special case of Markov-chain Monte Carlo (MCMC) simulation[27] where the dimension K of the distribution are sampled alternately one at a time, conditioned on the values of all other dimensions[22].

we apply the Gibbs sampling algorithm that has been introduced in [26], [22] to solve the SSN-LDA model and reduce the computation requirement.



Subsequently, the update equation for the hidden variable can be derived [22]:


where n(.)i is the count that does not include the current assignment of ιi and recall that sip is the variable for social interaction profiles. For the sake of simplicity, we assume that the Dirichlet distribution is symmetric in deriving the above formula.

Finally, the update formula for ϕk,w and θm,k are as follows:



In co-authorship networks, the vertices represent researchers and the edges in the network represent the collaboration relation between researchers. In this section we evaluate SSNLDA model on co-authorship networks collected from two distinct areas: computer science(CiteSeer) and nanotechnology(NanoSCI). Note that no name disambiguation has been done on either dataset.

The size of the largest connected subnetwork of CiteSeer is 249866 while the size of the largest connected subnetwork in NanoSCI is 203762. In this paper, we are only interested in discovering community structures in the two largest subnetworks.

In this paper, we explore three different types of social interaction pro le representations for social networks, namely 01-SIP, 012-SIP, and k-SIP.

In the 01-SIP approach, an edge is drawn between a pair of scientists if they coauthored one or more articles. Collaborating multiple times does not make a difference in this model

In order to mitigate this problem, we propose a 012-SIP model which takes a node's neighbors' neighbors into consideration. ... In this model, we distinguish a node's direct neighbors from its neighbors' neighbors by giving different weights to them.

This section describes a K-SIP model where the weight information for an edge is defined as the times of the collaboration between the two authors.

And then, 10% of the original datasets is held out as test set and we run the Gibbs sampling process on the training set for i iteration. In particular, in generating the exemplary communities, we set the number of the communities as 50, the iteration times i as 1000. In perplexity computation, i is set as 300 in order to shorten the computation time. In both case, α is set as 1/K and β is set as 0.01, where K is the number of the communities.

Table III shows 6 exemplary communities from a 50-community solution for the CiteSeer dataset with social interaction profiles being created using 012-SIP representation. ...These exemplary communities give us some flavor on the communities that can be discovered by this approach. Specifically, we observe that some communities are “institution-based”, some others are “topic-based. ... This observation reveals the fact that researchers from same institution or with similar research interests tend to collaborate together more and build closer social ties.

Perplexity is is a common criterion for measuring the performance of statistical models in information theory. It indicates the uncertainty in predicting the occurrence of a particular social interaction given the parameter settings, and hence it reflects the ability of a model to generalize unseen data.

Perplexity PP is defined as

where wm is the social interaction profiles in the test set and

where n(v)m is the number of times term t has been observed in document m.

It shows that the perplexity value is high initially and decreases when the number of communities increases. In addition, the results show that the 012-SIP approach has lower perplexity value than the other two approaches.

Compactness of a community is measured through the average shortest distance among the top-ranked Nr researchers in this community. Short average distance indicates a compact community.

The t-test results show that the 012-SIP approach is significantly better the other two approaches for both datasets

The probabilities that can be derived from SSN-LDA model can be helpful in determining the importance and roles of community members. For instance, the importance of community members conditioned on the community variable ιj can be measured through the probability p(wi|ιj ), which can be easily derived from this model based on the learned ϕ.

SSN-LDA model provides an elegant way to measure the similarity of two communities by calculating the corresponding KL (Kullback-Leibler) distance and convert it to similarity measure. KL divergence is a distance measure for two distributions and the corresponding formula for calculating the distance between two communities ιi and ιj is:



This paper describes an LDA (latent Dirichlet Allocation)-based hierarchical Bayesian algorithm, namely SSN-LDA(Simple Social Network LDA). In SSN-LDA, communities are modeled as latent variables in the graphical models and defined as distributions over social actor space. The advantage of SSN-LDA is that it only requires topological information as input.

2013年11月16日 星期六

Mimno, D., & McCallum, A. (2007, June). Mining a digital library for influential authors. In Proceedings of the 7th ACM/IEEE-CS joint conference on Digital libraries (pp. 105-106). ACM.

Mimno, D., & McCallum, A. (2007, June). Mining a digital library for influential authors. In Proceedings of the 7th ACM/IEEE-CS joint conference on Digital libraries (pp. 105-106). ACM.

本研究利用下面的公式來找到一個領域中具有影響力的作者
在式(1)裡,q代表用來查詢領域的問句,a代表用領域內的任一位作者,Pr(a|q)便代表給定問句找到作者a的可能性。d則是領域內的文件。Ad對於文件d而言代表其作者的集合,I{a屬於Ad}表示a是否為文件d的作者,1/(|Ad|)I{a屬於Ad}是任何一位作者在文件d上的機率Pr(a|d)。Pr(d)是個別文件的影響力,本研究根據Chen et al. [2]所建議的方式,利用PageRank演算法進行測量,代表研究人員根據參考文獻閱讀到這篇文件的可能性。

本研究比較三種估算Pr(q|d)的方式:
第一種方式利用Dirichlet平滑化的語言模型(language model),以下面的公式來計算:

此處w是文件d中出現的某一個詞語,Nwd是這個詞語出現在這個文件中的次數,Nw則是這個詞語出現在語料所有文件的次數。Nd是文件所所有詞語出現次數總和,N表示語料中所有詞語的總共出現次數。μ是一個調整變數,本研究將它設為100。

第二種方式從LDA(latent Dirichlete allocation)的主題模型(topic model)中挑選一個最符合問句q的主題t,以Pr(t|d)代替Pr(q|d)。

第三種方式同樣使用LDA的主題模型,但以下面的公式來估算Pr(q|d):

在主題模型的方式中使用LDA演算法的原因是它可以有效解決同義詞(synonymy)及同形詞(polysemy)的問題,因而將它視為是一種問句擴展(query expansion)的方法。LDA將文件視為各種主題的混合(mixtures),並且將每一個主題視為是詞彙上各個詞語的比例分布,而其原理便是利用語料中在文件裡經常一起出現的詞語來估測主題上各詞語的分布機率Pr(w|t),以及給定文件時出現各種主題的機率Pr(t|d)。

本研究以資訊檢索領域作為分析案例,結果發現第三種方式在發現具有影響力的作者上得到較好的結果。

We present a probabilistic model that ranks authors based on their influence in particular areas of scientific research. This model combines several sources of information: citation information between documents as represented by PageRank scores, authorship data gathered through automatic information extraction, and the words in paper abstracts.

Authors are coreferenced using Machine Learning methods. Previous studies of author influence in digital library collections have been hampered by ambiguities in authorship (for example, Newman [4] groups authors by first initial and last name). Finally, the link structure of the collection is identified by extracting and disambiguating the references from papers.

We measure the influence of individual documents using the PageRank algorithm. Chen et al. [2] demonstrate the use of PageRank on research literature, using references in place of hyperlinks. PageRank can be thought of as modeling a researcher who moves from paper to paper in the document collection. At each paper, the researcher either follows a randomly chosen reference from the current paper or, with probability , chooses a random paper from the collection. The PageRank of a given paper can be interpreted as the probability that the researcher will be reading that paper at any given moment. Since the PageRank is a probability distribution over all documents in the collection, we use it as the probability of a given document, Pr(d).

For the probability of authors given documents we use a uniform distribution, dividing the weight of a document evenly between its authors. The probability of an author given a document is Pr(a|d) = 1/(|Ad|), where Ad is the set of authors in paper d.

Using these elements we can construct a distribution over authors for a particular query,



where I{a屬於Ad} indicates whether a is listed as an author for a given paper.

For the component of the model that depends on the words in documents, Pr(q|d), we compare three statistical models. The first is based on a language model with Dirichlet smoothing. The second and third are based on a statistical topic model, using a single topic and a weighted sum of topics, respectively.

For the language model with Dirichlet smoothing, the probability of a query given a document is



where μ = 100, Nd is the number of words in document d, Nwd is the number of times word w appears in document d, Nw is the number of times word w appears in the corpus, and N is the total number of tokens in the corpus.

For the topic model we use Latent Dirichlet Allocation (LDA) [1]. LDA models documents as mixtures of "topics", which are probability distributions over the vocabulary of the corpus. Topic models are useful in handling synonymy (multiple words with similar meanings) and polysemy (words with multiple meanings), because they assign words to topics based on the context of the document. ... A trained topic model produces an estimate of the probability of a word given a topic, Pr(w|t), and the probability of a topic given a document, Pr(t|d).

In this application, the topic model can be considered a sort of query expansion: documents that contain none of the query words may still contain words that commonly occur in the same contexts as the query words.

In the second model we select a single topic t that matches the query and substitute Pr(t|d) for Pr(q|d) in Equation 1.

In the third model we represent Pr(q|d) as a weighted sum over all topics:

The weighted topic model approach to expert finding appears to be better able to generalize beyond the specific query words, while retaining a focus on areas relevant to the query. We believe that such models are a promising direction in expert finding, and a good example of the usefulness of structured digital library collections.

2013年4月8日 星期一

Rorissa, A. and Yuan, X. (2012). Visualizing and mapping the intellectual structure of information retrieval. Information Processing and Management, 48, 120-135.

Rorissa, A. and Yuan, X. (2012). Visualizing and mapping the intellectual structure of information retrieval. Information Processing and Management, 48, 120-135.

network analysis

Chen (2006)說明一個學術領域或學科的知識基礎(intellectual base)與研究前沿(research front)的區別:研究前沿是一個專業(specialty)當前最先進的技術狀態(state of the art),由研究前沿引用所構成的部分則是它的知識基礎。通常在分析學術領域或學科的知識基礎時,大多利用期刊的引用資料,並且使用群集分析(cluster analysis)、多維分析(multidimensional analysis)與其他技術將引用資料的視覺化 (例如:Chen & Kuljis, 2003; Chen et al., 2010; Ding et al., 2000; McKechnie, Goodall, LajoiePaquette, & Julien, 2005; Tang, 2004; White & Griffith, 1981; White & McCain, 1998)。White (2003)則是使用尋徑網路(Pathfinder Networks)技術將White & McCain (1998)的資料繪製成圖書資訊學領域的科學映射圖。這些技術也都被應用於軟體工具的製作並且提供研究人員自由使用,例如CiteSpace。將引用資料等資訊繪製成科學映射圖的研究興趣提升可以歸納以下的原因:可使用的引用資料來源以及其他廣泛出現;許多提供視覺化與映射圖的電腦應用程式可以自由使用;對不斷增加數量的電子資料需要具有包容性與容易使用的管理與理解方法。

這篇研究利用2000到2009年間的資訊檢索領域的引用資料進行一系列的分析與視覺化,所利用的資料來源是Web of Science,共使用48,390筆書目紀錄,使用的視覺化工具是由Chen(2004a; 2004b; 2006)提出的CiteSpace (http://cluster.cis.drexel.edu/cchen/citespace/)。研究的項目與結果如下:

1) 作者合著網路(co-authorship network): 以32位發表10篇以上的較高生產力作者為分析對象,,其中一半來自於電腦科學,另一半則為資訊科學。只要兩位作者曾一起出現在一篇論文中便在他們的映射點間建立連結。網路建立好之後,測量每位作者映射點的中介性,找出中心作者。結果可以發現兩個作者數目超過八位的較大作者叢集,中心分別是 Jarvelin K和Chen HC,前者來自資訊科學,而後者則是來自電腦科學。本研究推論作者的生產力較高,同時也將有較多的合著對象。
2) 高被引的期刊與論文:在這個研究裡發現約有43%的高被引文獻是在1970到1989年間出版,顯示資訊檢索已是一個相當成熟的領域。另外,生產力較高的作者群與高被引文獻間的關連不明顯。引用最高的出版品與作者都是來自電腦科學,引用最高的前三名作者以及他們的引用次數佔全部引用數的比率分別是Salton (26.1%), Jansen (9.88%), and Baeza-Yates (8.38%)。

3)作者指定的關鍵詞(author-assigned keywords):在關鍵詞所形成的共現網路裡,除了information retrieval,中介性較高的關鍵詞還包括information seeking、information system、evaluation和user studies。另外,在網路圖上可以明顯看出這些詞語形成的四個主要的研究集群為1)使用者研究(user studies)、網路資訊檢索(Web information retrieval)、3)引用分析/科學計量學(citation analysis/scientometrics)、和4)資訊檢索系統評估(information retrieval system evaluation)。
4) 活躍的機構(active institutions):資訊檢索領域的主要研究機構大部分位於美國,並且絕大多數是學術機構或大學。利用作者的合著關係將這些研究機構的合作情形畫成網路圖,結果發現這個研究領域相當鼓勵跨機構和跨國的研究。

5) 來自於其他領域的想法(the import of ideas from other disciplines):從論文的引用關係發現,資訊檢索研究的想法主要來自以下的五個領域:電腦科學(computer science)、圖書資訊學(library and information science)、工程(engineering)、電信傳播(telecommunications)和管理(management),高達91.6%的引用來自這五個領域。
We analyzed citation data in the information retrieval subfield for the past decade (2000–2009) and presented the results in terms of co-authorship network, highly productive authors, highly cited journals and papers, author-assigned keywords, active institutions, and the import of ideas from other disciplines.
An indicator of the maturity of an area of inquiry is the growth in the number and quality of research publications (Van den Beselaar & Leydesdorff, 1996). Insight into the nature of a field can be gained by examining ‘‘the publications produced by its practitioners. To the extent that practitioners in the field publish the results of their investigations, this mode for assessing the state of a field can reflect with great specificity the content and problem orientations of the group. Of the many ways that publications can be analyzed and counted, perhaps the most revealing kind of data are the references cited by the practitioner group in their publications’’ (Small, 1981, p. 39).
A field, discipline, or other area of study can be broadly divided into an intellectual base and current research fronts (Chen, 2006). ‘‘If we define a research front as the state of the art of a specialty (i.e., a line of research), what is cited by the research front forms its intellectual base’’ (Chen, 2006, p. 360). Previous literature of a discipline cited in its current publications (i.e., its intellectual base) can inform us about current research fronts. It is through references to sources that authors make connections between concepts (Small, 1981). Collectively, such connections create a ‘‘representation of the cognitive structure of the research field’’ (Small, 1981, p. 39).
Although a number of studies have examined the literature of library and information science (e.g., Åström, 2007, 2010; Chen, Ibekwe-SanJuan, & Hou, 2010; Cronin & Meho, 2007; Donohue, 1972; Harter & Hooten, 1992; Harter, Nisonger, & Weng, 1993; Persson, 1994; Rice, 1990; White & McCain, 1998; Zhao & Strotmann, 2008a, 2008b), the nature of the literature concerning information retrieval has not been so thoroughly investigated (Ding, Chowdhury, & Foo, 2000; Ding, Yan, Frazho, & Caverlee, 2009; Ellis, Allen, & Wilson, 1999).
According to Chen (2006), the trends and patterns of scientific literatures provide researchers or communities of similar interests an overview of the related field(s) and relationships among the specific fields. More specifically, such information as the most influential articles or books, the evolvement of terms, noun phrases, keywords, the most reputed researchers, the connection between different institutions and countries over time can show trends and patterns that provide more overview.
The most prominent conclusions address the stable, multidisciplinary nature of the field. For instance, Persson (1994) found that, for a 5-year period (1986–1990), the intellectual base of the Journal of the American Society for Information Science (consisting of the most frequently co-cited authors) and its research front (consisting of articles sharing at least five cited authors) had similar maps (depicted in two-dimensional spaces). This finding highlights the stable nature of the topics explored by the information studies field.
Ding et al. (2000) conducted one of the earliest studies on the literature of information retrieval specifically. They analyzed co-citation data of 50 highly cited journals using multidimensional scaling and cluster analysis. They produced two-dimensional maps of the structure of the literature of information retrieval for an 11-year period (1987–1997). The visualizations revealed strong relationships between information retrieval and five other disciplines: computer science, physics, chemistry, psychology, and neuroscience. Their maps also show information retrieval as being part of both the computer science and LIS fields.
Tang (2004) identified the most common disciplines to which LIS exports ideas (based on the number of citations it received from the disciplines):
computer science, communication, education, management, business, and engineering. Another study looked at the export and import of ideas to and from LIS and found it to be an exporter of ideas (Cronin & Meho, 2008). That contrasts sharply with the state of the field 20 years ago when few researchers from other disciplines cited LIS literature (Cronin & Pearson, 1990).
Result
Analysis of authorship and co-authorship is critical to the understanding of scholarly communication and knowledge diffusion (Chen, 2006).
To show the extent of collaboration by the most productive researchers in our dataset, we used CiteSpace to create a coauthorship network map. Two researchers have co-authorship if they have co-authored at least one paper together. CiteSpace generates networks by measuring betweenness centrality scores. ‘‘In a network, the betweenness centrality of a node measures the extent to which the node plays a role in pulling the rest of the nodes in the network together. The higher the centrality of a node, the more strategically important the node is.’’ (Chen et al., 2009, p. 236).
Two of the author co-citation clusters are large enough to contain at least eight members/authors. Two highly productive authors (see Table 1) anchor the two clusters: Jarvelin K and Chen HC. ... These results point to a relation between collaboration frequency and the most productive authors. We are tempted to conclude that the more actively an author collaborates, the more productive she or he is. Further research is necessary to confirm this assertion.

Chen, Cribbin, Macredie, and Morar (2002) showed that visualization can be used to track the development of a scientific discipline and present the long-term process of its competing paradigms. They also assert that, among a discipline’s co-cited publications, the cluster consisting of the most highly cited publications may represent the discipline’s core or predominant paradigm.
A product of CiteSpace, Fig. 2 displays a document co-citation network, generated from the collective citing behavior in our information retrieval dataset. The network is composed of 121 reference nodes and 1163 co-citation links.

Author-assigned keywords can reveal specific focus areas of research in a field.
As suggested by Table 4, ‘‘information retrieval’’ is located near the center of the cloud of keywords. Fig. 4 also indicates that other related terms have taken on a central role in the subfield: ‘‘information seeking,’’ ‘‘information system,’’ ‘‘evaluation,’’ and ‘‘user studies.’’ This highlights the rising emphasis on user-centered system design and retrieval, as well as the importance of user studies in the evaluation of IR systems. Information retrieval research is stronger today because it has increasingly focused on user-centered design. Current user studies research is more about ‘‘users’ interaction with information retrieval systems than about user information behavior in general’’ (Zhao & Strotmann, 2008a, p. 2077).
Fig. 4 also suggests that the information retrieval subfield has its own special areas of inquiry. Four main clusters can be discerned on the visualization map, centered around user studies, Web information retrieval, citation analysis/scientometrics, and information retrieval system evaluation.
Fig. 5 maps the collaboration between the top 20 institutions of information retrieval authors in our dataset. ... These diverse groupings indicate that the information retrieval subfield encourages collaboration across institutions and countries.
Information retrieval researchers in our dataset cite primarily computer science and library and information science publications (see Table 6). Those two fields account for 82.79% of the citations. ... Apart from LIS and computer science, the third, fourth, and fifth other disciplines from which information retrieval imports ideas are engineering, telecommunications, and management, respectively. In fact, 91.6% of the citations by information retrieval authors whose articles were published between 2000 and 2009 were to these five disciplines.