顯示具有 network analysis 標籤的文章。 顯示所有文章
顯示具有 network analysis 標籤的文章。 顯示所有文章

2017年12月7日 星期四

Nichols, L. G. (2014). A topic model approach to measuring interdisciplinarity at the National Science Foundation. Scientometrics, 100(3), 741-754.

Nichols, L. G. (2014). A topic model approach to measuring interdisciplinarity at the National Science Foundation. Scientometrics100(3), 741-754.

跨學科研究(IDR)是指由整合多個學科的理論、技術、資料與工具來解決單一學科無法解決的問題。在IDR的測量時,一般假定所有科學之間是一個連貫的學科結構(a coherent disciplinary structure),並且IDR的表現正是本質上模糊了學科之間的界限(Wagner et al., 2011)識別和測量IDR需要在一個研究計畫中評估多個學科的存在和整合,並且包括評估科學的投入,產出和過程(Wagner et al., 2011)。

測量跨學科性(interdisciplinary)的主要挑戰來自確定和界定構成IDR的不同學科。識別、理解和測量IDR需要對引導出研究方法、理論和結論的知識基礎(intellectual bases)進行解析和特徵化。研究人員已經嘗試了多種方法,但測量跨學科性及其隨著時間的動態仍然是一項艱鉅的工作。質性方法通常利用參與觀察、訪談和調查來描述多學科研究人員團隊中的過程和關係,並評估學科整合的程度(參見Masse et al, 2008; Stokols et al, 2003)。量化方法則通常仰賴於文獻計量學和網路分析技術(參見Porter和Rafols 2009; Rafols和Meyer 2008; Leydesdorff 2007),檢驗論文參考文獻列表中出現學科的引用分析是最常用的方法之一(Wagner et al., 2011)。許多有關測量IDR的文獻計量學文獻都側重於研究科學或出版物的產出(Wagner et al., 2011)。

本研究則是使用美國國家科學基金會(National Science Foundation, NSF)獎勵資料庫(award databse)的給獎建議與獎勵,從範圍更廣泛的人力、投入和過程等方面來測量IDR,描述IDR的互動與整合。

Gerrish和Blei(2010)有關測量學術影響(scholarly impact)的研究,比較了傳統的引文分析和基於語言的主題模型方法。他們發現,雖然這兩種方法在整體學術影響方面有一致的結果,但是基於語言的方法通常能確認在質量上有不同的有影響力的文章。Wang等(2011)結合LDA主題模型與網絡分析技術,開發幫助研究人員評估龐大且迅速增長的生物醫學文獻,以確定化學物質、基因和對藥物發現重要的疾病之間的顯著關聯的工具

本文利用NSF主題模型和NSF的體制結構(institutional structure),探討測量NSF獎勵組合中IDR的新方法。NSF主題模型(Newman et al., 2011)幫助NSF的工作人員與科學界更加了解NSF基金組合的內涵與脈絡,同時也提供文件學科內容(disciplinary content )的新評估方式2000年到2011年間由NSF頒發的獎項約170,000利用這些文件訓練NSF主題模型,共1000個主題,在扣除一些僅包含語言中常用的停字詞所組成的主題之後,共923個,並且依據主題對文件上的關連性,對每個獎項指定一到四個主題,主題的次序代表它們的關連性高低。

本研究利用前述運用主題模型方法產出的獎勵的指定主題,評估SBE(Social, Behavioral, and Economic Sciences)部門管理獎項跨學科

本研究利用NSF所屬的各部門代表學科,MPS (Mathematics and Physical Science) 因為包含多個彼此分離的學科,所以再細分第一步先對923個主題,利用所有170,000個獎項的指定結果以及獎項所屬NSF部門歸類。計算每個部門管理的獎項中每個主題出現的頻率,將主題指定給出現頻率最高的學科。如果有某一個主題高頻率地出現在多個部門或是被分配到非研究或是跨學科的部門,此時則進行個別檢視,根據主題描述將其指定給一個學科或是標示為「非學科特定」。非學科特定的主題例如,t3的假設(Hypothesis)、t60的儀器(Instrumentation)、t738的創業(Entrepreneurship)和t889的研究生(Graduate Students)。根據獎勵和主題在所有部門的分布統計,除了生物學(Biology)和地球科學(Geosciences)以外,其他部門兩者間的分布相當類似。生物學擁有10%的獎勵,但卻有18%的主題被指定給生物學,其原因可能是因為生物學包含多個次學科,而且各自使用相當專殊化的語言來描述他們的科學。反之,地球科學主管NSF23%的獎勵,卻只有6%的主題,其原因可能包括地球科學具有比較狹小的學科範圍、比較仰賴共同語言或者較強的跨學科連結。

本研究針SBE(Social, Behavioral, and Economic Sciences)部門管理的獎項進行跨學科性評估,因此在指定主題對應的主要部門後,再進一步針對SBE在2000到2011年間共有14,225個獎項,通過它們上面出現的主題所屬的學科數量以及主題在獎項上的出現排序,計算它們的跨學科性。如果獎勵包含主題的學科有一個被指定為SBE上其他的學科,則將該獎勵視為「內部跨學科性」(internal interdisciplinarity);若是該獎勵中只要有一個主題屬於其他部門或MPS的學科,則視為「外部跨學科性」(external interdisciplinarity),否則便是無跨學科性。除了上述簡單的三元化數量分析外,本研究也利用Stirling’s (2007)的多樣性指標(diversity index)評估每個獎項的跨學科性

此外,為了比較不同組合的科際整合性,本研究特別挑選6個核心計畫(core programs)進行分析: 社會學(Sociology)、政治學(Political Science)、經濟學(Economics)、地理空間學(Geography and Spatial Sciences, GSS)、決策、風險與管理科學(Decision Risk and Management Science, DRMS)和知覺、行動和認知科學(Perception, Action, Cognition, PAC),分析每個計畫內獎項組合的跨學科與Stirling多樣性指標,並且利用Sci2 Team (2009)的Science of Science Toolkit製作每個計畫的共現網路圖(co-occurrence network diagrams),對於從跨學科的各面向(數量、平衡與差異性)解釋和描述了不同類型的跨學科互動情形

研究結果發現,根據簡單的三元化數量分析,89%的SBE獎項是屬於跨學科研究,外部跨學科性和內部跨學科性分別占55%與34%,如果加上獎項的金額做為加權的話,有93%是跨學科研究,其中外部跨學科性高達74%,而內部跨學科性則是19%。其原因是獲得高額的獎勵大多是外部跨學科研究(約占79%),因為這些研究通常是需要較大成本與跨部門研究團隊的大型計畫



研究結果也顯示每年各類型(內部性、外部性)的跨學科研究數量和學科組合相當穩定,雖然各年之間主題分布有差異。而各年Stirling多樣性指標的平均值則是穩定而小幅成長,其原因可能是由於這些計畫大多為多年性的延續計畫。


六個核心計畫的跨學科研究獎勵比例與它們的平均Stirling多樣性指標有很大的差異,以結果來看,GSS、DRMS和PAC在獎勵比例較Economics和Political Science為大,平均Stirling多樣性指標也有同樣的結果,然而Sociology雖然跨學科研究的獎勵比例較大,然而它的Stirling多樣性指標卻較小。根據Stirling多樣性指標的計算方式,推斷造成Sociology在這項指標上較小的原因可能是在Sociology計畫內雖然許多學科都有跨學科研究的關係,但這些學科大多是SBE內的學科,只有少數SBE外的學科。此外,令一個可能的原因是Sociology內獎勵上使用的語言較一般,許多術語也常出現在其他學科中,所以僅僅只有兩個主題被歸類在這個學科,大部分相關的主題都被指定為非特定的SBE,因此在計算上獎勵在Sociology本身上的比例較小,而非特定的SBE的比例較大
以計畫的主題共現網路圖來分析,網路圖上節點代表該計畫內出現的各學科,節點大小表示相對應學科所占獎勵數量比例,節點間的連接線則代表兩個學科曾至少共同出現在一個獎勵,亦即它們之間曾有科際整合的記錄,線的粗細代表它們共同的獎勵數量比例。圖形上節點的數量表示對應計畫內學科的種類數量(variety),平均相連程度和網路密度則可以用來測量學科間的互動程度。例如,DRMS和Economics的網路上各有23個不同的學科,但是DRMS的平均相連程度和網路密度都比Economics來得大,分別是0.345 vs. 0.252和9.74 vs. 7.13,其原因是DRMS計畫內的獎勵大多由3到4個學科組成,而Economics的獎勵則只有2個學科。

2015年12月18日 星期五

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

近年來由於可以取得大量而多樣的文本資料和採用文本處理演算法等原因,研究人員對文本視覺化(text visualization)與視覺性的文本解析(visual text analytics)的研究興趣增加。本研究針對文本視覺化技術提出一個互動的視覺調查(visual survey)。並且利用此次調查的資料,分析文本視覺化的現況,比較研究使用的各種分析與視覺化技術,以及分析有關研究者的資訊,以提供搜尋相關研究、探索次領域(subfield)以及獲得研究趨勢的洞察等目的

本研究採納前人的研究,將文本視覺化技術,以分析任務(analytic tasks)、視覺化任務(visualization tasks)、資料領域(data domain)以及資料來源(data source)、資料性質(data property)、視覺化的維度(visualization dimensionality)、視覺化的呈現(visualization representation)、視覺化的排列方式(visualization alignment)等面向,建立分類架構(taxonomy)。


分析任務是指使用者採用文本視覺化技術預期達到的主要目的,這些分類包括:
1. 文本摘要 (Text Summarization) / 主題分析 (Topic Analysis) / 實體抽取 (Entity Extraction)
2. 言談分析 (Discourse Analysis):文本或對話轉錄(conversation transcript)裡流動的語言學分析。
3. 情感分析 (Sentiment Analysis)
4. 事件分析 (Event Analysis)
5. 趨勢分析 (Trend Analysis) / 樣式分析 (Pattern Analysis)
6. 詞法/語法分析 (Lexical / Syntactical Analysis)
7. 關係/連結分析 (Relation / Connection Analysis)
8. 翻譯/文本比對分析 (Translation / Text Alignment Analysis)

視覺化任務則是由文本視覺化技術所支援的較基層呈現與互動任務,包括:

1. 自動凸顯/建議興趣區 (Region of Interest)
2. 群集 (Clustering) / 分類 (Classification / Categorization)
3. 比較 (Comparison)
4. 概觀 (Overview)
5. 監視 (Monitoring)
6. 瀏覽 (Navigation) / 探索 (Exploration)
7. 對於不確定的對策 (Uncertainty Tackling)

資料領域,包括

1. 線上社交媒體 (Online Social media)
2. 通訊 (Communication)
3. 專利 (Patents)
4. 評論 (Reviews) / 病歷 (Medical Records)
5. 文學作品 (Literature) / 詩 (Poems)
6. 科學文章 (Scientific Articles) / 論文 (Papers)
7. 社論媒體 (Editorial Media)

資料來源有單一文件 (Document) [33]、語料庫 (Corpora) [25]以及 串流文本 (Streams) [19];特殊的資料性質包括地理空間 (Geospatial) [11]、時間序列 (Timeseries) [14] 以及網路 (Networks) [6];視覺化的再現包括下列項目:折線圖 (Line Plot) / 河流圖 (River) [9, 18]、像素 (Pixel) / 面積 (Area) / 矩陣 (Matrix) [13, 7, 4]、節點-連結 (Node-Link) [32]、雲 (Clouds) / 銀河 (Galaxies) [1, 3]、地圖 (Maps) [34]、文本 (Text) [26]與形符 (Glyph) / 圖標 (Icon) [28, 10];排列則包括了輻射狀 (Radial) [35]、線性 (Linear) / 平行線 (Parallel) [8] 以及測標依賴 (Metric-dependent) [22]。

本研究指出有超過一半(56%)的文本視覺化利用主題模型(topic modeling)技術,資料來源方面大多數支援語料庫(70%),並且許多支援時間相關的資料(43%),而視覺再現方面主題以二維(2-D)為主,僅有極少數的研究以三維(3-D)的方式呈現,約占所有研究的4%。

文本視覺化的前五位主要作者為Daniel A. Keim (17 筆)、Shixia Liu (12 筆)、Christian Rohrdantz (9 筆)、Daniela Oelke (7 筆)和 Huamin Qu (7 筆)。將作者依據他們的合著關係建立研究者合作網路圖後,觀察網路圖的相連成分,可以發現大部分是獨立的小群體,最大的成分上共有106位作者,並且在這個成分上的兩個主要集群為University of Konstanz和Microsoft Research Asia等兩個研究團隊,Daniel A. Keim 和 Shixia Liu分別為集群的中心,並且他們二位也是網路圖上中介中心性最高的節點。雖然在本研究蒐集的資料上,這兩位作者之間並沒有直接的合作關係,但他們都曾與中介中心性第三高的兩位作者Dongning Luo 和 Jing Yang合作。


In this paper, we present an interactive visual survey of text visualization techniques that can be used for the purposes of search for related work, introduction to the subfield and gaining insight into research trends.

The interest for text visualization and visual text analytics has been increasing for the last ten years. The reasons for this development are manifold, but for sure the availability of large amounts of heterogeneous text data (caused by the popularity of online social media) and the adoption of text processing algorithms (e.g., for topic modeling) by the InfoVis and Visual Analytics communities are two possible explanations.




Analytic Tasks
these items are critical to the main analysis goals that users expect to achieve when employing a text visualization technique.

1. Text Summarization / Topic Analysis / Entity Extraction

2. Discourse Analysis
the linguistic analysis of the flow of text or conversation transcript.

3. Sentiment Analysis
for techniques related to the analysis of sentiment, opinion, and affection.

4. Event Analysis
deal with the extraction of events from the text data or involve visualization of text in some different manner

5. Trend Analysis / Pattern Analysis
both automated trend analysis and manual investigation directed at discovering patterns in the textual data.

6. Lexical / Syntactical Analysis

7. Relation / Connection Analysis

8. Translation / Text Alignment Analysis

Visualization Tasks
lower-level representation and interaction tasks that are supported by the text visualization techniques.

1. Region of Interest
the automatic highlighting/suggestion of data items/regions that could be of interest to the user for more detailed investigation

2. Clustering / Classification / Categorization

3. Comparison

4. Overview
both techniques that provide “the big picture” by displaying a significant portion of the data set as well as techniques which use special aggregated representations to provide overview while reducing the visual complexity

5. Monitoring

6. Navigation / Exploration

7. Uncertainty Tackling


Domain

1. Online Social media

2. Communication

3. Patents

4. Reviews / (Medical) Records

5. Literature / Poems

6. Scientific Articles / Papers

7. Editorial Media

Data sources include the following self-evident items: Document [33], Corpora [25], and Streams [19].

The special data properties include Geospatial [11], Timeseries [14], and Networks [6].

Representation includes the following items: Line Plot / River [9, 18], Pixel / Area / Matrix [13, 7, 4], Node-Link [32], Clouds / Galaxies [1, 3], Maps [34], Text [26], and Glyph / Icon [28, 10].

Alignment, i.e., layout, includes Radial [35], Linear / Parallel [8], and Metric-dependent [22].

As displayed in the table, our proposed taxonomy includes most of the categories except for two: we believe that the underlying data representation (e.g., bag-of-words vs. language model [30] or whole text vs. partial text [24]) is more relevant to the underlying computational methods than to observable visualization techniques.

And the same naturally holds for data processing methods (e.g., the specification of involved MDS methods [2]) that are partially covered by other categories in our taxonomy, for instance, the analytic task of topic analysis implies the usage of corresponding computational methods.

Using the data collected for the survey, we have been able to analyze the general state of the text visualization field, to compare the usage of various analysis and visualization techniques (with regard to our taxonomy), and to analyze the information about researchers in this field.

According to our current set of entries, the trend for rapid increase of text visualization techniques started around 2007.

With regard to category statistics (cf. Fig. 4), there is an obvious interest for tasks related to topic modeling (56% of all entries).

The majority of the techniques support corpora as data sources (70% of all entries), and a lot of them support time-dependent data (43% of all entries).

Another result—which is probably expected—is that only less than 4% of all entries use 3-dimensional visual representations.

We have also taken a look at the authorship statistics for the current data set. The top five authors with regard to number of techniques are Daniel A. Keim (17 entries), Shixia Liu (12 entries), Christian Rohrdantz (9 entries), Daniela Oelke (7 entries), and Huamin Qu (7 entries).

As seen in Fig. 5, the majority of author nodes are included into isolated connected components of small sizes (less than 10 nodes) while there is a big connected component with 106 nodes present in the graph.

The two major clusters in that component represent the research groups from the University of Konstanz and Microsoft Research Asia with Daniel A. Keim and Shixia Liu as cluster center nodes.

Shixia Liu and Daniel A. Keim happen to have the 1st and the 2nd largest betweenness values in the graph, respectively. While these two researchers have no direct collaboration with regard to our data set, they both have collaborated with Dongning Luo and Jing Yang who both share the 3rd largest betweenness value.

2015年4月15日 星期三

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V. (2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V.(2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

由於一般認為將領域之間的關係表示為圖形,通過考慮這些關係的可能性能夠提供許多資訊,不論對新進人員或專家皆有助於理解與分析,因此對這方面方法與工具的需求逐漸提高。過去的研究大多以期刊為分析單位,產生所有科學研究領域的科學映射圖。例如Leydesdorff (2004a, 2004b)使用雙重連結成分(biconnected components)的圖形分析演算法,將JCR 2001的科學研究進行分類。Boyack, Klavans, and Börner (2005)則應用了8種不同的期刊相似性測量7121種SCI和SSCI期刊,並採用VxOrd產生科學映射圖。Samoylenko, Chao, Liu, and Chen (2006)建構科學期刊的最小生成樹(minimum spanning trees),他們使用的資料是SCI 1994到2001的資料。本研究提出一個將ISI (Institute of Scientific Information)類別繪製成科學映射圖的方法,這個方法利用根據類別間的共被引資訊建構類別間的連結,以尋徑網路(PathfinderNetwork)縮減不重要的連結,然後以Kamada-Kawai方法決定節點在圖上的布局(layout),最後利用因素分析(factor analysis)進行結構確認。本研究和先前的研究都是針對類別利用共被引資訊呈現科學映射圖。以類別為分析單位在代表上足夠明確,並且比起較小的單位,這種方式對非專家使用者(nonexpert user)較具有資訊且使用者友善。Moya-Anegón et al. (2004)針對西班牙科學研究領域的視覺化,Moya-Anegón et al. (2005)則進一步利用科學映射圖比較英國、法國和西班牙三個國家的科學研究領域。本研究依循Börner, Chen, and Boyack (2003)提出的知識領域映射流程。使用的資料為7585種ISI期刊,ISI的類別共有219個,但扣除多學科科學後(Multidisciplinary Sciences),採用的類別共218個。利用共被引計算期刊相似性的方式為

Cc(ij)為期刊i和期刊j共被引次數,c(i)和c(j)則分別是期刊i和期刊j被引用次數。然後以尋徑網路和Kamada-Kawai方法繪製網路圖,經過尋徑網路處理後,有較多連結的節點具有較重要的地位。而尋徑網路是一種以型態為主的方法,與以群集為主的因素分析彼此間可以互補,因素分析可以識別、界定與定名科學映射圖上呈現的主題區域,而尋徑網路則負責讓使主題區域更加明顯,將類別分組成束,並顯示連接不同顯著類別的路徑,以及總體的型態結構。。最後總計共分析出35個因素,通過陡坡考驗(scree test)則有16個。科學映射圖上的類別可以分為三個群集:醫學與地球科學、基礎與實驗科學以及社會科學。

This study proposes a new methodology that allows for the generation of scientograms of major scientific domains, constructed on the basis of cocitation of Institute of Scientific Information categories, and pruned using PathfinderNetwork, with a layout determined by algorithms of the spring-embedder type (Kamada–Kawai), then corroborated structurally by factor analysis.

We present the complete scientogram of the world for the Year 2002.

This need arises from the general conviction that an image or graphic representation of a domain favors and facilitates its comprehension and analysis, regardless of who is on the receiving end of the depiction and whether a newcomer or an expert.

Science maps can be very useful for navigating around in scientific literature and for the representation of its spatial relations (Garfield, 1986). They are optimal means of representing the spatial distribution of the areas of research while also offering additional information through the possibility of contemplating these relationships (Small & Garfield, 1985).

From a general viewpoint, science maps reflect the relationships between and among disciplines; but the positioning of their tags clues us into semantic connections while also serving as an index to comprehend why certain nodes or fields are connected with others.

Moreover, these large-scale maps of science show which special fields are most productively involved in research—providing a glimpse of changes in the panorama—and which particular individuals, publications, institutions, regions, or countries are the most prominent ones (Garfield, 1994).

It is a tool in that it allows the generation of maps, and a method in that it facilitates the analysis of domains, by showing the structure and relations of the inherent elements represented. In a nutshell, scientography is a holistic tool for expressing the discourse of the scientific community it aspires to represent, reflecting the intellectual consensus of researchers on the basis of their own citations of scientific literature.

In Moya-Anegón et al. (2004), we ventured forth with a historic evolution of scientific maps from their origin to the present, and proposed ISI-JCR category cocitation for the representation of major scientific domains. Its utility was demonstrated by a visualization of the scientific domain of geographical Spain for the Year 2000.

Since then, other works related with the visualization of great scientific domains have appeared; however, all use journals as the unit of analysis, with the exception of a study based on the cocitation of categories (Moya-Anegón et al., 2005), comparatively focusing on three geographic domains (England, France, and Spain).

In contrast, Leydesdorff (2004a, 2004b) classified world science using the graph-analytical algorithm of biconnected components in combination with JCR 2001.

Boyack, Klavans, and Börner (2005) applied eight alternative measures of journal similarity to a dataset of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes, to show the first global map of science using the force-directed graph layout tool VxOrd.

Samoylenko Chao, Liu, and Chen (2006) proposed an approach through the construction of minimum spanning trees of scientific journals, using the Science Citation Index from 1994 to 2001.

In processing and depicting the scientific structure of great domains, we further developed a methodology that follows the flow of knowledge domains and their mapping as proposed by Börner, Chen, and Boyack (2003).

Because ISI assigns each journal to one or more subject categories, to designate a subject matter (i.e., ISI category) for each document, we also downloaded the Journal Citation Report (JCR; Thomson Corporation, 2005a), in both its Science and Social Sciences editions, for 2002.

The downloaded records were exported to a relational database that reflects the structured information of the documents. This new repository contained nearly 1 million (N = 901,493) source documents: articles, biographical items, book reviews, corrections, editorial materials, letters, meeting abstracts, news items, and reviews that had been published in 7,585 ISI journals (N = 5,876 + 1,709). These were classified in a total of 219 categories, altogether citing 25,682,754 published documents.

As informational units, they are, in themselves, sufficiently explicit to be used in the representation of all disciplines that make up science in general. These categories, in combination with the adequate techniques for the reduction of space and the representation of the information to construct scientograms of science or of major scientific domains, prove much more informative and user friendly for quick comprehension and handling by nonexpert users than those obtained by the cocitation of smaller units of cocitation.

For these reasons, we used the 219 categories of the JCR 2002 as units of measure, with the exception of “Multidisciplinary Sciences.” ... The maximum number of categories with which we worked, then, was 218.

In light of our previous experience (Moya-Anegón et al., 2004, 2005), we use cocitation as the similarity measure to quantify the relationship existing between each one of the JCR categories.

Therefore, after a number of trials, we arrived at the conclusion that using tools of Network Analysis, the best visualizations are those obtained through raw data cocitation as the unit of measure. Yet, it also was necessary to reduce the number of coincident cocitations to enhance pruning algorithm yield. Therefore, to those raw data values we added the standardized cocitation value. In this way, we could work with raw data cocitation while also differentiating the similarity values between categories with equal cocitation frequencies. The key was a simple modification of the equation for the standardization of the degree of citation proposed by Salton and Bergmark:




where CM is cocitation measure, Cc is cocitation frequency, c is citation, and i and j are categories.

Over the history of the visualization of scientific information, very different techniques have been used to reduce n-dimensional space. Either alone or in conjunction with others, the most common are multidimensional scaling, clustering, factor analysis, self-organizing maps, and PathfinderNetworks (PFNET).

In our opinion, PFNET with pruning parameters r = ∞, and q = n − 1 is the prime option for eliminating less significant relationships while preserving and highlighting the most essential ones, and capturing the underlying intellectual structure in a economical way.

Although PFNET has been used in the fields of Bibliometrics, Informetrics, and Scientometrics since 1990 (Fowler & Dearhold, 1990), its introduction in citation was due to the hand of Chen (1998, 1999), who introduced a new form of organizing, visualizing, and accessing information. The end effect is the pruning of all paths except those with the single highest (or tied highest) cocitation counts between categories (White, 2001).

The spring embedder type is most widely used in the area of documentation, and specifically in domain visualization. Spring embedders begin by assigning coordinates to the nodes in such a way that the final graph will be pleasing to the eye (Eades, 1984). Two major extensions to the algorithm proposed by Eades (1984) have been developed by Kamada and Kawai (1989) and Fruchterman and Reingold (1991).

While Brandenburg, Himsolt, and Rohrer (1995) did not detect any single predominating algorithm, most of the scientific community goes with the Kamada–Kawai algorithm. The reasons upheld are its behavior in the case of local minima, its capacity to minimize differences with respect to theoretical distances in the entire graph, good computation times, and the fact that it subsumes multidimensional scaling when the technique of Kruskal and Wish (1978) is applied.

We can effortlessly see which are the most important nodes in terms of the number of their connections and, in turn, which points act as intermediaries with other lines, as hubs or forking points.

Whereas factor analysis is a clustering-oriented procedure, PFNET is topology oriented. Yet, they are extremely valuable as complements in the detection of the structure of a scientific domain.

Thus, factor analysis is responsible for identifying, delimiting, and denominating the great thematic areas reflected in the scientogram.

Meanwhile, PFNET is in charge of making the subject areas more visible, grouping their categories into bunches, and showing the paths that connect the different prominent categories, and finally, the overall topology of the domain.

Factor analysis identifies 35 factors in the cocitation matrix of 218 × 218 categories of world science 2002. Through the scree test we extracted 16, which we tagged using the previously explained method; these accumulate 70.2% of the variance (Table 1)

The number of categories included in at least one factor is 195. Twenty-three were not included in any factor (Table 2), and 25 belonged to two factors simultaneously (Table 5).

That is, a category or thematic area occupying a central position in the scientogram will have a more general or universal nature in the domain as a consequence of the number of sources it shares with the rest, contributing more to scientific development than those with a less central position.

The more peripheral the situation of a category or subject area, the more exclusive its nature, and the fewer the sources it will appear to share with other categories; accordingly, the lesser its contribution to the development of knowledge through scientific publications.

An intermediary position favors the interconnection of other categories or thematic areas. 

This broad interpretation of our scientograms not only explains the patterns of cocitation that characterize a domain but also foments an intuitive way for specialists and nonexperts to arrive at a practical explanation of the workings of PFNET (Chen & Carr, 1999).

From a macrostructural point of view, we can distinguish three major zones.

In the center is what we could call Medical and Earth Sciences, consisting of Biomedicine, Psychology, Etiology, Animal Biology & Ecology, Health Care & Service, Orthopedics, Earth & Space Science, and Agriculture & Soil Sciences.

To the right, we can see some other basic and experimental sciences: Materials Sciences & Physics, Applied; Engineering; Computer Science & Telecommunications; Nuclear Physics & Particles & Fields; and Chemistry.

To the left is the neighborhood of the social sciences, with Applied Mathematics, Business, Law, and Economy, and Humanities.

On one hand, it offers domain analysts the possibility of seeing the most essential connections between categories of given domain.

On the other hand, it allows us to see how these categories are grouped in major thematic areas, and how they are interrelated in a logical order of explicit sequences.

2015年3月30日 星期一

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

確認科學領域的專業(specialties)本質是資訊科學的一項基本挑戰 (Morris & Van der Veer Martens, 2008; Tabah, 1999) 。由於1)可取用的書目資料來源愈來愈普及;2)網路上愈來愈多可提供分析與視覺化的電腦軟體工具;3)從多元來源而大量的資料吸收的要求愈來愈劇烈等原因,因此有愈來愈多的相關研究。共被引分析是對科學進行量化分析最常用的方法之一,特別是作者共被引分析 (author cocitation analysis, ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b)以及文件共被引分析 (document cocitation analysis, DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985)。作者共被引分析的目的在透過被相關文獻一起引用的作者群集,確認領域裡的專業。重要的作者共被引分析研究包括White & McCain (1998),這個研究以1972到1995年間12種資訊科學相關期刊的120位高被引作者進行作者共被引分析,研究結果發現當時的資訊科學分為兩個基本上彼此獨立的陣營:資訊檢索(information retrieval)與文獻(literature)。Zhao and Strotmann (2008a, 2008b) 以1996-2005年的資訊科學相關期刊資料重新進行了相同的研究,他們的結果發現了5個主要的專業:使用者研究(user studies)、引用分析(citation analysis)、實驗型檢索(experimental retrieval)、網路計量學 (Webometrics)以及知識領域的視覺化(visualization of knowledge domains),其中新興的兩個專業:網路計量學和知識領域的視覺化連繫了引用分析以及實驗型檢索,而使用者研究則是此時最大的專業。Aström (2007) 則是使用文件共被引分析的例子,他們分析了1990到2004年的21種圖書資訊學期刊,利用多維尺度法(multidimensional scaling, MDS)產生結果,他們的結果與White & McCain (1998)的研究類似,整個領域可分為兩個陣營,不過Aström (2007)的結果將稱為資訊尋求與檢索(information seeking and retrieval),而不是資訊檢索。

不管是作者共被引分析或是文件共被引分析其步驟大致如下:
1) 檢索引用資料。
2) 建構參考文件或作者共同被引用的矩陣。
3) 將共被引矩陣表示成節點與連結的圖(node-and-link graph)或是多維尺度法的組態(configuration),並且可以利用尋路網路(Pathfinder network scaling)或最小生成樹(minimum spanning tree)裁減連結。
4) 利用群集、社群發現(community finding)、因素分析(factor analysis)、主成分分析(principle component analysis)或者隱含語意索引(latent semantic indexing)等各種演算法確認專業。例如Morris & Van der Veer Martens (2008)、 Persson (1994)、 Tabah (1999)、 White & Griffith (1982)以及Janssens, Leta, Glänzel, and De Moor (2006)。
5) 根據群集成員間共同的主題(themes),解釋共被引群集的性質。通常需要豐富的領域知識,而且是一個花費大量時間與認知需求(cognitively demanding)的工作。

本研究對於作者共被引以及文件共被引形成的群集進行結構與動態的描述與解釋,分析的資料為1996到2008年間的12種資訊科學(information science)領域相關期刊,共計10853筆書目紀錄,引用的參考文獻為129060筆,引用次數為206180,而參考文獻的作者共有58711位。本研究以餘弦(cosine)測量作者或文件之間的關連大小,做為節點間的連結,建立網路;然後計算從原先網路導出的Laplacian矩陣(Laplacian matrices)的特徵向量(eigenvectors)找出群集。這種利用標準線性代數的頻譜群集(spectral cluster)演算法,較其他的群集演算法更有效率,而且因為不需要假設群集的形式,所以更有彈性與強健。標註群集方面則是利用引用文獻論文的詞語與摘要句,詞語包括題名與摘要中出現的名詞片語與索引詞(index terms),利用 tf*idf (Salton, Yang, & Wong, 1975)、對數似然比(log-likelihood ratio, LLR)測試 (Dunning, 1993)以及相互資訊(mutual information, MI)等三種資訊做為判斷的參考。摘要句則是從題名與摘要尋找最有代表性的句子,例如以Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007)對句子進行排序。




A multiple-perspective cocitation analysis method is introduced for characterizing and interpreting the structure and dynamics of cocitation clusters.

The generic method is applied to a three-part analysis of the field of information science as defined by 12 journals published between 1996 and 2008: (a) a comparative author cocitation analysis (ACA), (b) a progressive ACA of a time series of cocitation networks, and (c) a progressive document cocitation analysis (DCA).

Identifying the nature of specialties in a scientific field is a fundamental challenge for information science (Morris & Van der Veer Martens, 2008; Tabah, 1999).

The growing interest in mapping and visualizing the structure and dynamics of specialties is because of a number of reasons:
1. Widely accessible bibliographic data sources such as the Web of Science, Scopus, and Google Scholar (Bar-Ilan, 2008; Meho & Yang,2007) as well as domain-specific repositories such as ADS (http://www.adsabs.harvard.edu/) and arXiv (http://arxiv.org/).
2. Freely available computer programs and Web-based general-purpose visualization and analysis tools such as ManyEyes (http://manyeyes.alphaworks.ibm.com/) and Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/; Batagelj & Mrvar, 1998), special-purpose citation analysis tools such as CiteSpace (http://cluster.cis.drexel.edu/&u0007E;cchen/citespace/; Chen, 2004; Chen, 2006), and social network analysis such as UCINET (http://www.analytictech.com/ucinet6/ucinet.htm).
3. Intensified challenges for digesting the vast volume of data from multiple sources (e.g., e-Science, Digging into Data (http://www.diggingintodata.org/), cyber-enabled discovery, SciSIP; Lane, 2009).

Cocitation studies are among the most commonly used methods in quantitative studies of science, especially including author cocitation analysis (ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b) and document cocitation analysis (DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985).

For instance, once cocitation clusters are identified, assigning the most meaningful labels for these clusters is currently a challenging task because any representative labels of clusters must characterize not only what clusters appear to represent, but also the salient and unique reasons for their formation.

The new procedure reduces analysts' cognitive burden by automatically characterizing the nature of a cocitation cluster in terms of (a) salient noun phrases extracted from titles, abstracts, and index terms of citing articles and (b) representative sentences as summarizations of clusters.

ACA aims to identify underlying specialties in a field in terms of groups of authors who were cited together in relevant literature.

White & McCain (1998) presented a comprehensive view of information science based on 12 journals in library and information science across a 24-year span (1972–1995). It analyzed cocitation patterns of 120 most-cited authors with factor analysis and multidimensional scaling. The authors drew upon their extensive knowledge of the field and offered an insightful interpretation of 12 specialties identified in terms of 12 factors. The most well-known finding of the study is that information science at the time consisted of two essentially independent camps, namely, the information retrieval camp and the literature camp, including citation analysis, bibliometrics, and scientometrics.

Zhao and Strotmann (2008a, 2008b) followed up White and McCain's study using the same set of 12 journals and the same number of 120 cited authors in an updated time frame of 1996-2005. ... Zhao and Strotmann (2008b) found five major specialties and manually labeled them as user studies, citation analysis, experimental retrieval, Webometrics, and visualization of knowledge domains. In contrast to the findings of (White & McCain, 1998), experimental retrieval and citation analysis retained their fundamental roles in the field, and the user studies specialty became the largest specialty. Webometrics and visualization of knowledge domains appeared to make connections between the retrieval camp and the citation analysis camp.

A DCA by Aström (2007) studied papers published between 1990 and 2004 in 21 library and information science journals. Results were depicted in multidimensional scaling (MDS) maps. Aström's study also identified the two-camp structure found by (White & McCain, 1998). On the other hand, Aström found an information seeking and retrieval camp, instead of the information retrieval camp as in (White and McCain).

Although manually labeling a cocitation cluster can be a very rewarding process of learning about the underlying specialty and result in insightful and easy to understand labels, it requires a substantial level of domain knowledge and it tends to be time-consuming and cognitively demanding because of the synthetic work required over a diverse range of individual publications.

Traditionally, researchers often identify the nature of a cocitation cluster based on common themes among its members. ... The emphasis on common areas is a practical strategy; otherwise, comprehensively identifying the nature of a specialty can be too complex to handle manually.

Many researchers have studied the structural and dynamic properties of specialties in information science in terms of clusters, multivariate factors, and principle components (Morris & Van der Veer Martens, 2008; Persson, 1994; Tabah, 1999; White & Griffith, 1982).

A recent study of information science (Ibekwe-SanJuan, 2009) mapped the structure of information science at the term level using a text analysis system TermWatch and a network visualization system Pajek, but it did not address structural patterns of cited references.

Researchers also studied the structure of information science qualitatively, especially with direct inputs from domain experts. For example, Zins conducted a Critical Delphi study of information science, involving 57 leading information scientists from 16 countries (Zins, 2007a, 2007b, 2007c, 2007d).

Janssens, Leta, Glänzel, and De Moor (2006) studied the full-text of 938 publications in five library and information science journals with latent semantic analysis (LSA; Deerwester, Dumais, Landauer, Furnas, & Harshman, 1990) and agglomerative clustering. They found an optimal 6-cluster solution in terms of a local maximum of the mean silhouette coefficients (Rousseeuw, 1987) and a stability diagram (Ben-Hur, Elisseeff, & Guyon, 2002). Their clusters were labeled with single-word terms selected by tf*idf (p. 1625), which are not as informative as multiword terms for cluster labels.

Klavans, Persson, and Boyack (2009) recently raised the question of the true number of specialties in information science. They suspected that the number is much more than the 11 or 12 as reported in ACA studies such as (White & McCain, 1998) and (Zhao & Strotmann, 2008a, 2008b), but significantly fewer than the 72 reported in their own study, which is also based on the 12 journals between 2001 and 2005.

The 12-journal Information Science dataset, retrieved from the Web of Science, contains 10,853 unique bibliographic records, written by 8,408 unique authors from 6,553 institutions and 89 countries. These articles cited 129,060 unique references for a total of 206,180 times. They cited 58,711 unique authors and 58,796 unique sources.

The traditional procedure of cocitation analysis for both DCA and ACA comprises the following steps:
1. Retrieve citation data from sources such as the Science Citation Index (SCI), Social Science Citation Index (SSCI), Scopus, and Google Scholar.
2. Construct a matrix of cocited references (DCA) or authors (ACA).
3. Represent the cocitation matrix as a node-and-link graph or as a multidimensional scaling (MDS) configuration with possible link pruning using Pathfinder network scaling or minimum spanning tree algorithms.
4. Identify specialties in terms of cocitation clusters, multivariate factors, principle components, or dimensions of a latent semantic space using a variety of algorithms for clustering, community finding, factor analysis, principle component analysis, or latent semantic indexing.
5. Interpret the nature of cocitation clusters.

The interpretation step is the weakest link. It is time-consuming and cognitively demanding, requiring a substantial level of domain knowledge and synthesizing skills. In addition, much of attention routinely focuses on cocitation clusters per se, but the role of citing articles that are responsible for the formation of such cocitation clusters may not be always investigated as an integral part of a specialty.

Our new method extends and enhances traditional cocitation methods in two ways: (a) by integrating structural and content analysis components sequentially into the new procedure and (b) by facilitating analytic tasks and interpretation with automatic cluster labeling and summarization functions. The new procedure is highlighted in yellow in Figure 2, including clustering, automatic labeling, summarization, and latent semantic models of the citing space (Deerwester et al., 1990).

Our new procedure adopts several structural and temporal metrics of cocitation networks and subsequently generated clusters.

Structural metrics include betweenness centrality, modularity, and silhouette.

Temporal and hybrid metrics include citation burstness and novelty

The betweenness centrality metric is defined for each node in a network. It measure the extent to which the node is in the middle of a path that connects other nodes in the network (Brandes, 2001; Freeman, 1977). High betweenness centrality values identify potentially revolutionary scientific publications (Chen, 2005) as well as gatekeepers in social networks.

In the context of this study, the modularity Q measures the extent to which a network can be divided into independent blocks, i.e., modules (Newman, 2006; Shibata, Kajikawa, Taked, & Matsushima, 2008).

The silhouette metric (Rousseeuw, 1987) is useful in estimating the uncertainty involved in identifying the nature of a cluster.

Burst detection determines whether a given frequency function has statistically significant fluctuations during a short time interval within the overall time period.

Sigma is introduced in (Chen, et al., 2009a) as a measure of scientific novelty. ... In this study, Sigma is defined as (centrality + 1)burstness such that the brokerage mechanism plays more prominent role than the rate of recognition by peers.

We adopt a hard clustering approach such that a cocitation network is partitioned to a number of nonoverlapping clusters.

In this article, cocitation similarities between items i and j are measured in terms of cosine coefficients.

A good partition of a network would group strongly connected nodes together and assign loosely connected ones to different clusters. This idea can be formulated as an optimization problem in terms of a cut function defined over a partition of a network. Technical details are given in relevant literature (Luxburg, 2006; Ng, Jordan, & Weiss, 2002; Shi & Malik, 2000).

Spectral clustering is an efficient and generic clustering method (Luxburg, 2006; Ng et al., 2002; Shi & Malik, 2000). It has roots in spectral graph theory. Spectral clustering algorithms identify clusters based on eigenvectors of Laplacian matrices derived from the original network.

Spectral clustering has several desirable features compared to traditional algorithms such as k-means and single linkage (Luxburg, 2006):
 • It is more flexible and robust because it does not make any assumptions on the forms of the clusters,
• it makes use of standard linear algebra methods to solve clustering problems, and
• it is often more efficient than traditional clustering algorithms.

Candidates of cluster labels are selected from noun phrases and index terms of citing articles of each cluster. These term are ranked by three different algorithms. In particular, noun phrases are extracted from titles and abstracts of citing articles. The three term ranking algorithms are tf*idf (Salton, Yang, & Wong, 1975), log-likelihood ratio (LLR) tests (Dunning, 1993), and mutual information (MI).

Each cocitation cluster is summarized by a list of sentences selected from the abstracts of articles that cite at least one member of the cluster.

In this study, sentences are ranked by Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007). Given a set S of N sentences, let M be the square matrix that for each pair of sentences gives the number of nominal words in common (nouns and adjectives).

In this study, summarization sentences were also ranked by two new functions gtf and gftidf , which are further simplified approximations of the energy function E.

The ACA and DCA studies described in this article were conducted using the CiteSpace system (Chen, 2004; Chen, 2006). CiteSpace is a freely available Java application for visualizing and analyzing emerging trends and changes in scientific literature.

CiteSpace supports a unique type of cocitation network analysis—progressive network analysis—based on a time slicing strategy and then synthesizing a series of individual network snapshots defined on consecutive time slices. Progressive network analysis particularly focuses on nodes that play critical roles in the evolution of a network over time. Such critical nodes are candidates of intellectual turning points.

In summary, (a) spectral clustering and factor analysis identified about the same number of specialties, but they appeared to reveal different aspects of cocitation structures and (b) cluster labels chosen from citers of a cluster tend to be more specific terms than those chosen by human experts.

We found the comparison with the study of Zhao and Strotmann very valuable. It offered us an opportunity to compare the analysis conducted by human experts to the interpretation cues provided by our automatic labeling and summarization methods.

Spectral clustering for the purpose of network decomposition is exclusive in nature although in reality it is often sensible to allow overlapping clusters because of multiple roles individual entities may play.

Spectral clustering of cocitation networks tends to generate distinct clusters with high precision, whereas human experts tend to aggregate entities into broadly defined clusters.

In conclusion, the new cocitation analysis procedure has the following advantages over the traditional one:
• It can be consistently used for both DCA and ACA.
• It uses more flexible and efficient spectral clustering to identify cocitation clusters.
• It characterizes clusters with candidate labels selected by multiple ranking algorithms from the citers of these clusters and reveals the nature of a cluster in terms of how it has been cited.
• It provides metrics such as modularity and silhouette as quality indicators of clustering to aid interpretation tasks.
• It provides integrated and interactive visualizations for exploratory analysis.

Modularity and silhouette metrics provide useful quality indicators of clustering and network decomposition.

2014年8月10日 星期日

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers. Scientometrics, 94(3), 1161-1173.

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers.Scientometrics, 94(3), 1161-1173.

network analysis

本研究以發表場域-作者-耦合(Venue-Author-Coupling,VAC)、期刊共被引分析(journal co-citation analysis)、主題分析(topic analysis)和連結編輯委員會成員(interlocking editorial board membership)等四個面向分析資訊科學與圖書館學的期刊網絡。這個研究分析的期刊範圍為2008年 JCR (Journal Citation Report)資訊科學與圖書館學分類的58種期刊,在2005到2009年間的出版資料。分析資料的相關數據如Table 1:


本研究利用VAC代表期刊的生產者(producers)的相似性,根據每一對期刊間相同的作者數量測量它們的接近程度,其原理建立在作者會選擇主題或社會性相似(thematically or socially similar)的期刊發表。期刊共被引分析(McCain, 1991)計算每一對期刊被共同引用的次數,本研究用來測量作品(artifacts)間的相似程度。本研究以修改自LDA模型(Blei et al. 2003)的ACT(Author-Conference-Topic)模型(Tang et al., 2008)透過關鍵詞(keywords)在主題上的分布以及主題在作者及發表場域(期刊)上的分布,本研究以餘弦(cosine)測量評估期刊之間的相似程度。連結編輯委員會成員則是編輯委員會上的共同成員數測量期刊間的相似程度。兩種期刊間共同的成員愈多,代表這兩種期刊在認知上或是社會性上愈相似。

根據上面的四種期刊間的相似程度所得到的結果,除了進行階層式集群分析(hierarchical cluster analysis)之外,也用來建立網絡,以Kamada-Kawaii 法呈現網絡的型態。分析得到的四種網絡並且以二次指派程序(Quadratic Assignment Procedure) (Lawler 1963)比較網絡之間可能的相關性(correlation)。

VAC方法得到的期刊網絡如下
四個集群分別為MIS(黃)、IS(藍)、LS(綠)以及專門性期刊(紅)。其中的MIS期刊集群與其他的集群相當分離。IS與LS距離較近。相較於其他三個期刊集群,專門性期刊彼此間的連結較弱。
期刊共被引分析所得的網絡如下:

主題模型產生的五個主題如Table 2
五個主題在網絡上的分布如下圖
MIS(黃)仍然與其他集群較為分離,但與健康和傳播(communication)等專門性期刊的距離較近。IS(藍)和LS(粉紅)的位置與VAC和期刊共被引分析的網絡上有所不同。圖書館服務與實務(綠)與專門性期刊和LS很接近。

在利用連結編輯委員會成員的期刊網絡上,有10種期刊沒有和其它期刊有共同編輯委員。其餘的集群分為四群。以傳播研究相關的期刊是新增加的集群(綠)。

四個網絡的QAP結果如Table 3

總結以上,在JCR的資訊科學與圖書館學分類下約略可以將期刊分為四個集群:MIS、IS、LS和傳播相關的期刊。MIS相較來說較為獨立。另外,QAP的結果可以看到編輯委員會成員的結果與期刊共被引分析有很高的相關性,其原因可能是由於擔任編輯委員的研究人員往往有較好的學術成就,被引用的機會較高。編輯委員會成員與VAC有較高的相關性,其原因也可能是編輯委員有較高的生產力。運用多種面向的分析可以較全面地了解整個學術傳播網絡。

Fifty-eight journals from the Information Science and Library Science category in the 2008 Journal Citation Report were studied and the network proximity of these journals based on Venue-Author-Coupling (producer), journal co-citation analysis (artifact), topic analysis (concept) and interlocking editorial board membership (gatekeeper) was measured. The resulting networks were examined for potential correlation using the Quadratic Assignment Procedure.

The VAC approach is used to represent the producers in this dataset. This approach measures journal proximity based on the number of authors shared by each journal pair. The VAC approach is based on the idea that an author’s choice of publication venue reflects similarity judgments authors are likely to choose venues that are thematically or socially similar.

Artifacts are measured by means of journal co-citation. This measure, introduced by McCain (1991), refers to the appearance of two journals in the same reference list of an article. The more frequently two journals appear in the same reference lists, the greater the similarity between the two journals. The journal co-citation approach measures journal proximity by the frequency with which each journal pair is co-cited by the same articles.

Topic modeling is used to capture concepts. ... The technique adopted here, the author-conference-topic (ACT) model (Tang et al., 2008), extends the LDA model by considering the author and publishing venue of the articles. LDA was developed originally as a topic modeling technique concerning the probability distribution of keywords for topics, and is particularly helpful with the ‘‘classification, novelty detection, summarization, and similarity and relevance judgment’’ of large-scale data (Blei et al. 2003, p. 993). ... This model extends the idea of LDA by taking into account the authors and publishing venues, and estimates not only the distribution of words on topics, but also the distribution of authors and venues on the topics modeled. ... Here, the outcome of the ACT model is the probability distribution of each author and each journal over topics, and the journal proximity is calculated using the cosine similarity of the journals.

The interlocking editorship approach, employed by Ni and Ding (2010), measures journal proximity based on common editorial board membership. The number of editorial board members that two journals share can be viewed as an indicator of journal similarity. ... Thus, it can be expected that if two journals have scholars in common on their editorial boards, these two journals have some degree of similarity, either cognitively or socially.

The journals were clustered using a hierarchical clustering technique with squared Euclidean distance and Ward’s method. Each journal clustering was displayed as a network (Kamada-Kawaii layout); each node (journal) was colored according to the hierarchical clustering result with the size of a
node proportional to its centrality (either degree or closeness).

Additionally, a comparison of journal proximity results was conducted using the Quadratic Assignment Procedure (QAP). QAP is commonly used in social network analysis as a means of investigating correlations between two networks. ... (Lawler 1963).

2014年2月28日 星期五

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of American Society for Information Science and Technology, 57(3), 359-377.

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature.  Journal of American Society for Information Science and Technology, 57(3), 359-377.

information visualization

本研究提出一個整合研究專業(specialty)的研究前沿(research front)以及其引用的知識基礎(intellectual base)的視覺化介面。本論文定義研究前沿為研究專業上一組急遽出現的概念(concepts)與研究議題(research issues);研究前沿的知識基礎則是包含這些概念與研究議題的論文引用或者共同被引用的論文。在針對某一個專業進行其研究前沿與知識基礎進行視覺化時,首先蒐集專業相關的論文,從這些論文抽取代表研究前沿的詞語,並以論文所引用或共被引的論文做為專業的知識基礎,建立分別代表研究前沿的詞語和知識基礎的論文的二方網路(bipartite networks)以同時呈現研究前沿的相關概念與研究議題以及知識基礎的論文。在建立起來的網路上透過詞語和論文形成的叢集可以發現重要的研究前沿和知識基礎,藉由詞語呈現叢集的概念與研究議題更能有效地表達研究前沿的意涵,並且如果加上論文的發表時間來分析,可以從急遽出現在較多論文的相關詞語找出發展中的研究前沿。此外,對於網路進行中介中心性(centrality of betweenness)分析可以發現研究前沿間具有樞紐地位的論文,並且透過Pathfinder演算法可以發現論文間的主要關連。
A specialty is conceptualized and visualized as a time-variant duality between two fundamental concepts in information science: research fronts and intellectual bases.
A research front is defined as an emergent and transient grouping of concepts and underlying research issues.
The intellectual base of a research front is its citation and co-citation footprint in scientific literature— an evolving network of scientific publications cited by
research-front concepts.
The concept of a research front was originally introduced by Price (1965) to characterize the transient nature of a research field. Price observed what he called the immediacy factor: There seems to be a tendency for scientists to cite the most recently published articles. In a given field, a research front refers to the body of articles that scientists actively cite.
A specialty can be conceptualized as a time-variant mapping from its research front to its intellectual base.
Typical questions regarding a research front may include:
How did it get started? What is the state of the art? What are the critical paths in its evolution?
To address such questions, we need to detect and analyze emerging trends and abrupt changes associated with a research front over time. We also need to identify the focus of a research front at a particular time in the context of its intellectual base, to reveal significant intellectual turning points as a research front evolves, and to discover the interconnections between different research fronts.
Braam, Moed, and Raan (1991) defined a specialty as “focused attention by a number of scientific researchers to a set of related research problems and concepts” (p. 252). They studied the continuity and stability of a specialty in terms of the similarity between co-citation clusters across consecutive years. The similarity between two co-citation clusters is determined by comparing aggregated word profiles of the clusters.
In part, this is because we define a research front differently to emphasize emerging trends and abrupt changes as the defining features of a research front. A research front is the domain of a time-variant mapping, and its intellectual base is the co-domain of the mapping.
Griffith et al. (1974) found that between-cluster co-citation links tend to be weaker than within-cluster co-citation links. ... To understand how specialties and different thematic trends interact with each other, it is essential to study the nature of long-range, between-cluster links and understand why articles in different specialties were connected.
Labeling clusters is concerned with the clarity and interpretability of co-citation clusters. The standard approach relies on word profiles derived from articles citing a cluster of co-cited articles. ... Word-profile approaches have drawbacks. First, word profiles may not converge to a focused message. Analysts and users will make a substantial amount of sense-making efforts to synthesize a diverse range of word profiles. Second, cluster labels based on aggregating word profiles tend to be too broad to be useful. In practice, many users would be interested in not only the most commonly used terms but also terms that can lead to profound changes. Terms associated with an emerging trend could be overshadowed by a broader and more persistent theme.
In CiteSpace II, a current research front is identified based on such burst terms extracted from titles, abstracts, descriptors, and identifiers of bibliographic records. These terms are subsequently used as labels of clusters in heterogeneous networks of terms and articles.
CiteSpace II makes it easier for users to identify pivotal points. In addition to inspecting salient visual attributes, the user easily can see nodes with high betweenness centrality (Freeman, 1979).
The procedure of using CiteSpace II is described in the following steps, 
(1) Identify a knowledge domain using the broadest possible term.
(2) Data collection
(3) Extract research front terms: CiteSpace II first collects n-grams, or terms, from titles, abstracts, descriptors, and identifiers of citing articles in a dataset. The present study used single words or phrases of up to four words. ... Research-front terms are determined by the sharp growth rate of their frequencies.
(4) Time slicing
(5) Threshold selection
(6) Pruning and merging: Pathfinder network scaling is the default option in CiteSpace II for network pruning (Chen, 2004; Schvaneveldt, 1990).
(7) Layout
(8) Visual inspection
(9) Verify pivotal points
We demonstrate the new features of CiteSpace with case studies of two research fields: mass-extinction research (1981–2003) and terrorism research (1990–2003).
Mass-extinction research (1981–2003).
The input data for CiteSpace II were retrieved from citation index databases via the Web of Science based on a topic search for articles published between 1981 and 2003 on mass extinction. The scope of the search included four topic fields in each bibliographic record: title, abstract, descriptors, and identifiers. The search was limited to articles in English only.
The resultant dataset contains a total of 771 records.
A total of 333 research-front terms were detected from the four topic fields of these records.
Terrorism research (1990–2003).
The terrorism research (1990–2003) dataset consists of 1,776 records resulted from a topic search on terrorism in the Web of Science.
A total of 1,108 research-front terms were found.
The fully integrated representation of research fronts and intellectual bases in the same network visualization has three practical advantages.
First, using surged topical terms rather than the most frequently occurring title words is particularly suitable for detecting emerging trends and abrupt changes. In visualized networks, research-front terms are explicitly linked to intellectual-base articles. This design presents a compact representation of the duality between a research front and its intellectual base.
Second, research-front terms naturally lend themselves to be used as labels of specialties.
Third, it overcomes a common drawback of word-profile-based labeling approaches. Aggregated word profiles may not converge to an intrinsic focus. Terms selected based on sudden increased popularity measures are particularly suitable to characterize a current research front.
The Pathfinder algorithm extracts the most salient patterns from a network, but it does not scale well. CiteSpace II implements a concurrent version of the algorithm. The concurrent Pathfinder algorithm has substantially optimized the network scaling module, although it still took 6,000 seconds to process 14 networks and merge them into a 1,704-node network.
In conclusion, the new features introduced to CiteSpaceII for detecting and visualizing emerging trends and abrupt changes in a field of research have produced promising and encouraging results. The major findings are that
• the surge of interest is an informative indicator for a new research front;
• using heterogeneous networks of terms and articles provides a comprehensive representation of the dynamics of a specialty;
• research-front terms are informative cluster labels;
• citation tree-ring visualizations are visually appealing and semantically interpretable;
• betweenness centrality metrics identify semantically valid pivotal points.