顯示具有 information visualization software 標籤的文章。 顯示所有文章
顯示具有 information visualization software 標籤的文章。 顯示所有文章

2015年4月21日 星期二

Tseng, Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

Tseng,  Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

近幾十年來,發展出許多科學計量分析技術,包括為了群集(clustering)書目資料所需的各種相似度(similarity)計算技術,如共被引(co-citation)、書目耦合(bibliographic coupling)與詞語共現分析(co-word analysis),這些技術的比較分析可參見Yan and Ding (2012)。並且有很多可以在網路上自由下載使用的軟體工具製作並包裝這些技術,提供科學計量分析應用,知名的軟體工具如CiteSpace (Chen 2006, Chen et al. 2010)、Sci2 Tool (Sci2 Team 2009)、VOSviewer (Van Eck and Waltman 2010)、BibExcel (Persson 2009)及Sitkis (Schildt and Mattsson 2006),這部分的分析則可參見Cobo et al. (2011)。本研究包含兩個部分:提出包含一系列利用書目計量資訊進行群集與映射(mapping)技術的科學計量分析軟體工具集 CATAR,並且將此工具集應用於圖書資訊學(library and information science, LIS)領域後,希望能夠利用期刊群集的結果,確認與分析次領域,以及建議適合研究評估(research evaluation)用途的LIS期刊集合。

Åström (2002)從領域概念的視覺化研究獲得一個結論:期刊的選擇確實影響研究領域如何被知覺與定義,也就是研究領域的界定(delineation)與期刊的選擇有密切關係。已經有許多的研究對圖書資訊學進行次領域界定,而這些研究大多參考ISI的JCR主題分類中與圖書資訊學最相關的類別IS&LS(Information Science and Library Science)。IS&LS類別下並不只包含圖書資訊學的相關期刊,這個類別涵蓋兩個密切相關的領域資訊科學(Information Science)和圖書館學(Library Science),此一範圍與圖書資訊學有些微不同。根據Leydesdorff (2008),JCR主題分類以期刊的題名、引用模式(citation patterns)等等做為標準進行分類,但是這個分類結果與從資料庫本身的引用資料所產生的網路上的主要成分(principal components)得到的分類結果並不十分相符。因此次領域界定研究大多經過人為的挑選做為分析資料的期刊,並沒有完整收錄IS&LS主題下的所有期刊。

進行次領域界定時常使用的技術包括:利用共被引分析比較一對項目,利用凝聚式階層群集(agglomerative hierarchical clustering, AHC)將項目分群產生樹狀圖(dendrogram),利用多維尺度(multi-dimensional scaling, MDS)產生視覺化的二維或三維映射圖。若干重要的研究如:Åström (2002)從圖書資訊學重要期刊中選取1135篇出版在1998到2000年的文章,利用BibExcel軟體工具進行作者共被引(author co-citation)以及關鍵詞共現分析,並產生MDS映射圖,52位高被引作者的共被引產生三個群集:"硬"資訊檢索(hard information retrieval)、"軟"資訊檢索(soft information retrieval)以及書目計量學(bibliometrics),47個較常出現的關鍵詞則分為圖書館學(library science,LS)、資訊檢索(information retrieval,IR)及書目計量學。Åström (2002)認為作者共被引分析沒有出現圖書館學的原因可能與圖書館學研究的出版管道有關,如果引用的資料像是書籍或地區期刊沒有出現在JCR,圖書館學作者便無法出現在引用為基礎的排名上。Åström (2007)對55種在JCR 2003主題類別下的期刊,選擇21種圖書資訊學相關期刊的13605篇文章進行文件共被引分析,在從1990到2004年的三個時段發現圖書資訊學可分為資訊計量學(informetrics)和資訊搜尋與檢索(information seeking and retrieval)兩個穩定的次領域,而隨著全球資訊網的普及,網路計量學(webometrics)在兩個次領域上都成為主要的研究議題。Jassen et al. (2006) 對2002到2004年五種圖書資訊學相關期刊的938篇文章,應用一系列的全文分析技術以及MDS和AHC,將938篇文章分為六個群集:兩個群集與書目計量學有關、一個群集為IR、一個包含一般議題、另兩個較小但愈來愈重要的群集分別是網路計量學和專利分析(patent analysis)。Moya-Anegon et al. (2006)從24種較有影響力的期刊中選擇17種期刊,排除將資訊科學(information science, IS)應用到特定技術或知識領域(例如:醫學、地理學、電訊傳播等),從17種期刊引用的參考文獻,對77位最常被引用的作者和73篇最常被引用的期刊進行共被引分析,映射使用的技術包括MDS和AHC以及自組織映射圖(self-organizing map)。作者共被引分析的結果產生六個次領域:科學計量學、引用分析、書目計量學、"軟"(認知導向)資訊檢索、"硬"(演算法導向)資訊檢索以及傳播理論(communication theory)。而期刊共被引分析的結果則有四個群集:IS、LS、科學研究(science studies)以及管理學(management)。在期刊共被引分析的科學研究大致上可以對應為作者共被引分析的科學計量學、引用分析、書目計量學,IS為"軟"資訊檢索和"硬"資訊檢索。如Åström (2002)同樣的原因,LS沒在作者共被引分析的結果當中。Waltman et al. (2011)以JASIST為種子,選擇與該期刊共被引較多的期刊,連JASIST共48種,進行期刊的書目耦合(bibliographic coupling)分析,並且利用VOSviewer呈現視覺化結果,共分為LS、IS以及科學計量學等3個次領域。Milojevic et al. (2011)使用詞語共現分析探討1998到2007年出版的16種期刊上的10344篇文章,16種期刊根據Nisonger and Davis (2005) 的研究所挑選,分析100個文章題名上最常出現的詞語,進行共現分析,並以AHC歸類,結果三個主要群集為LS、IS以及書目計量學/科學計量學。

Åström (2002)以關鍵詞的共現分析所得到的結果包括LS次領域,但作者共被引分析所得到的映射圖上並沒有產生這個次領域。Moya-Anegon et al. (2006)的期刊共被引分析與作者共被引分析也略有不同,期刊共被引分析的結果上有作者共被引分析沒有的LS和管理學兩個次領域,反之,作者共被引分析的結果上則可以發現期刊共被引分析沒有的傳播學理論(communication theory)。一般認為這和作者引用的行為有關,LS作者的引用次數大多沒有達到分析的門檻,因此無法在上述兩個研究的作者共被引分析結果上呈現。

Ni et al. (2012)從JCR的IS&LS類別下的61種期刊,排除3種非英語的期刊,將選取的58種期刊進行場域-作者耦合(venue-author coupling)、期刊共被引分析、詞語共現分析、期刊連結(journal interlocking)等四種分析。分析的結果再進行MDS與AHC分析,四種方式所得到一致的次領域包括:管理資訊系統(managment information systems, MIS)、IS、LS和特殊化群集(specialized clusters),並且在四種方法所得到MDS映射的圖形上都可以發現MIS與其他群集分離,Ni and Ding (2010)與Ni and Sugimoto (2011)建議JCR上的圖書資訊相關期刊應進行適當的重組。

本研究(Tseng and Tsai 2013)應用的資料範圍為2000到2004與2005到2009在Web of Science 的Journal Citation Report中 Information Science & Library Science (IS&LS)主題分類下的所有期刊,在前期(2000~2004年)共50種,後期(2005~2009年)共66種。本研究的分析程序採用Borner et al. (2003)整理的一般工作流程,步驟包括:1) 資料蒐集(data collection);2)文本分段(text segmentation);3)相似性計算(similarity computation);4)多階段群集(multi-stage clustering);5)群集標名(clustering labeling);6)視覺化(visualization);7)面向分析(facet analysis)。這些步驟中所需的技術都已經整合到軟體工具CATAR(Content Analysis Toolkit for Academic Research, http://web.ntnu.edu.tw/~samtseng/CATAR/)上。在計算文件間的相關性時,本研究以一種期刊做為一個文件,所有論文引用的期刊做為文件的特徵,然後利用Dice係數(Salton 1989)計算期刊相似性,例如兩種期刊X與Y,R(X)與R(Y)分別是它們引用的期刊,它們之間的相似性計算為Sim(X, Y) = 2 ∙ |R(X)∩R(Y)|/(|R(X)|+R(Y)|)。也就是利用書目耦合計算期刊之間的相似性。期刊的群集則是利用完全連接階層群集法(complete-linkage hierarchical clustering)。首先將每個文件視為一個群集,然後將一對最相似的群集合併起來,產生一個較大的群集,然後重複進行上面的步驟,而兩個群集的相似性定義為兩個群集間最小的文件相似性,如果相似性超過某個預先設定的閾值,便將兩個群集合併,一直到無法再產生合併為止。此外,本研究採用Silhouette指標(Ahlgren and Jarneving 2008; Rousseuw 1987; Jassen et al. 2006)。

此一研究的資料包含JCR的IS&LS主題下的期刊,分為2000-2004年與2005-2009年兩個時期,前一個時期包含50種期刊,9546筆論文資料;後一時期則有66種期刊,11471筆論文資料。從群集結果的樹狀圖(dendrogram)和MDS映射的結果顯示,IS&LS主題下的期刊在兩個時期都有IR、MIS、科學計量學、學術圖書館(academic library)、醫學圖書館(medical library)、館藏發展(collection development),以及開放取用(open access)和地區圖書館(regional library)兩個後期出現並且較小的群集。並且MIS群集的期刊在知識基礎(intellectual base)上與IS&LS主題的其他期刊分離,表示這群集下的期刊具有較特殊的引用模式。本研究以期刊的書目耦合進行分析,從期刊知識基礎(intellectual base)得到MIS群集與其他分離的研究結果,與Ni et al. (2012)利用期刊共被引分析、期刊連結、術語使用(terminology usage)和合著(co-authorship)研究等不同方法的研究結果相同,這也為許多探討圖書資訊學認知結構的研究認為不應將MIS相關期刊與其他期刊包含在ISI的同一個主題IS&LS下,在進行分析時需要排除MIS相關期刊提供了佐證(Larivière et al. 2012)。此外,並且以多樣性指標(diversity index)分析群集特性,揭露出某些次領域具有地區(regional)特性。

2014年2月28日 星期五

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of American Society for Information Science and Technology, 57(3), 359-377.

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature.  Journal of American Society for Information Science and Technology, 57(3), 359-377.

information visualization

本研究提出一個整合研究專業(specialty)的研究前沿(research front)以及其引用的知識基礎(intellectual base)的視覺化介面。本論文定義研究前沿為研究專業上一組急遽出現的概念(concepts)與研究議題(research issues);研究前沿的知識基礎則是包含這些概念與研究議題的論文引用或者共同被引用的論文。在針對某一個專業進行其研究前沿與知識基礎進行視覺化時,首先蒐集專業相關的論文,從這些論文抽取代表研究前沿的詞語,並以論文所引用或共被引的論文做為專業的知識基礎,建立分別代表研究前沿的詞語和知識基礎的論文的二方網路(bipartite networks)以同時呈現研究前沿的相關概念與研究議題以及知識基礎的論文。在建立起來的網路上透過詞語和論文形成的叢集可以發現重要的研究前沿和知識基礎,藉由詞語呈現叢集的概念與研究議題更能有效地表達研究前沿的意涵,並且如果加上論文的發表時間來分析,可以從急遽出現在較多論文的相關詞語找出發展中的研究前沿。此外,對於網路進行中介中心性(centrality of betweenness)分析可以發現研究前沿間具有樞紐地位的論文,並且透過Pathfinder演算法可以發現論文間的主要關連。
A specialty is conceptualized and visualized as a time-variant duality between two fundamental concepts in information science: research fronts and intellectual bases.
A research front is defined as an emergent and transient grouping of concepts and underlying research issues.
The intellectual base of a research front is its citation and co-citation footprint in scientific literature— an evolving network of scientific publications cited by
research-front concepts.
The concept of a research front was originally introduced by Price (1965) to characterize the transient nature of a research field. Price observed what he called the immediacy factor: There seems to be a tendency for scientists to cite the most recently published articles. In a given field, a research front refers to the body of articles that scientists actively cite.
A specialty can be conceptualized as a time-variant mapping from its research front to its intellectual base.
Typical questions regarding a research front may include:
How did it get started? What is the state of the art? What are the critical paths in its evolution?
To address such questions, we need to detect and analyze emerging trends and abrupt changes associated with a research front over time. We also need to identify the focus of a research front at a particular time in the context of its intellectual base, to reveal significant intellectual turning points as a research front evolves, and to discover the interconnections between different research fronts.
Braam, Moed, and Raan (1991) defined a specialty as “focused attention by a number of scientific researchers to a set of related research problems and concepts” (p. 252). They studied the continuity and stability of a specialty in terms of the similarity between co-citation clusters across consecutive years. The similarity between two co-citation clusters is determined by comparing aggregated word profiles of the clusters.
In part, this is because we define a research front differently to emphasize emerging trends and abrupt changes as the defining features of a research front. A research front is the domain of a time-variant mapping, and its intellectual base is the co-domain of the mapping.
Griffith et al. (1974) found that between-cluster co-citation links tend to be weaker than within-cluster co-citation links. ... To understand how specialties and different thematic trends interact with each other, it is essential to study the nature of long-range, between-cluster links and understand why articles in different specialties were connected.
Labeling clusters is concerned with the clarity and interpretability of co-citation clusters. The standard approach relies on word profiles derived from articles citing a cluster of co-cited articles. ... Word-profile approaches have drawbacks. First, word profiles may not converge to a focused message. Analysts and users will make a substantial amount of sense-making efforts to synthesize a diverse range of word profiles. Second, cluster labels based on aggregating word profiles tend to be too broad to be useful. In practice, many users would be interested in not only the most commonly used terms but also terms that can lead to profound changes. Terms associated with an emerging trend could be overshadowed by a broader and more persistent theme.
In CiteSpace II, a current research front is identified based on such burst terms extracted from titles, abstracts, descriptors, and identifiers of bibliographic records. These terms are subsequently used as labels of clusters in heterogeneous networks of terms and articles.
CiteSpace II makes it easier for users to identify pivotal points. In addition to inspecting salient visual attributes, the user easily can see nodes with high betweenness centrality (Freeman, 1979).
The procedure of using CiteSpace II is described in the following steps, 
(1) Identify a knowledge domain using the broadest possible term.
(2) Data collection
(3) Extract research front terms: CiteSpace II first collects n-grams, or terms, from titles, abstracts, descriptors, and identifiers of citing articles in a dataset. The present study used single words or phrases of up to four words. ... Research-front terms are determined by the sharp growth rate of their frequencies.
(4) Time slicing
(5) Threshold selection
(6) Pruning and merging: Pathfinder network scaling is the default option in CiteSpace II for network pruning (Chen, 2004; Schvaneveldt, 1990).
(7) Layout
(8) Visual inspection
(9) Verify pivotal points
We demonstrate the new features of CiteSpace with case studies of two research fields: mass-extinction research (1981–2003) and terrorism research (1990–2003).
Mass-extinction research (1981–2003).
The input data for CiteSpace II were retrieved from citation index databases via the Web of Science based on a topic search for articles published between 1981 and 2003 on mass extinction. The scope of the search included four topic fields in each bibliographic record: title, abstract, descriptors, and identifiers. The search was limited to articles in English only.
The resultant dataset contains a total of 771 records.
A total of 333 research-front terms were detected from the four topic fields of these records.
Terrorism research (1990–2003).
The terrorism research (1990–2003) dataset consists of 1,776 records resulted from a topic search on terrorism in the Web of Science.
A total of 1,108 research-front terms were found.
The fully integrated representation of research fronts and intellectual bases in the same network visualization has three practical advantages.
First, using surged topical terms rather than the most frequently occurring title words is particularly suitable for detecting emerging trends and abrupt changes. In visualized networks, research-front terms are explicitly linked to intellectual-base articles. This design presents a compact representation of the duality between a research front and its intellectual base.
Second, research-front terms naturally lend themselves to be used as labels of specialties.
Third, it overcomes a common drawback of word-profile-based labeling approaches. Aggregated word profiles may not converge to an intrinsic focus. Terms selected based on sudden increased popularity measures are particularly suitable to characterize a current research front.
The Pathfinder algorithm extracts the most salient patterns from a network, but it does not scale well. CiteSpace II implements a concurrent version of the algorithm. The concurrent Pathfinder algorithm has substantially optimized the network scaling module, although it still took 6,000 seconds to process 14 networks and merge them into a 1,704-node network.
In conclusion, the new features introduced to CiteSpaceII for detecting and visualizing emerging trends and abrupt changes in a field of research have produced promising and encouraging results. The major findings are that
• the surge of interest is an informative indicator for a new research front;
• using heterogeneous networks of terms and articles provides a comprehensive representation of the dynamics of a specialty;
• research-front terms are informative cluster labels;
• citation tree-ring visualizations are visually appealing and semantically interpretable;
• betweenness centrality metrics identify semantically valid pivotal points.

2014年1月24日 星期五

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

本研究探討在地理上的生產力遷移(shifts in productivity)是否發生在書目計量學(bibliometrics)、資訊計量學(informetrics)和科學計量學(scientometrics)等計量學(metrics)領域,也就是歐洲的貢獻明顯地成長,並且北美的貢獻相對來說有減少的情形。

有關計量學的研究,Hood and Wilson (2001)和Stock and Weber(2006)等研究都分析了這個領域的文獻成長情形。Hood and Wilson (2001)回顧了計量學領域的發展,並且比較bibliometrics、scientometrics和informetrics的相關文獻,發現bibliometrics還是在相關領域上使用最廣泛的詞語。Stock and Weber(2006)從觀察中確認這個領域從1980年後便持續地成長。Wolfram (2008)則發現在計量學領域中,北美的文獻有明顯地減少而歐洲則是急遽地增加的情形。

本研究利用bibliometrics、scientometrics、informetrics、cybermetrics、webometrics、citation analysis、link analysis和citation indexes做為檢索的問句,同時再加上Scientometrics和Journal of Informetrics兩種期刊的論文,從Web of Science資料庫中進行檢索。結果共檢索出4404筆論文資料。

在這些論文資料裡,共有75個國家。以地區來區分,歐洲在每個時段上具有最大的貢獻,不論是數量或所占比率都有成長,亞洲所佔的相對比例在22年間有很大的成長,北美雖然在數量上有成長,可是相對的比例呈現緩慢的下降。每個地區的作者會偏好在本身地區的期刊上發表,舉例而言,歐洲作者發表論文的前五個期刊中有四個歐洲期刊,南美也有類似的情形,但是亞洲的情形例外,前五個期刊中有四個是歐洲期刊,另一個則是北美的期刊。

自1990年代中期後,國家間的合作情形增加許多,之前國際合作的論文每年為1到19篇,2008年已大幅增加為96篇。美國是國際合作佔最多的國家,但以地區來說,歐洲平均每個國家的國際合作數為5.78篇論文,多於世界其他部分的4.47篇論文。

此外,歐洲則有許多具有國際合作經驗的機構,共有16所研究機構有國際合作經驗,北美則有8所,亞洲有1所。機構間的合作來說,在1987年每篇論文平均只有1.1個機構,但在2007年則增加為1.96。

本研究且利用MDS、VOSviewer和Pajek將這些論文上的國家與機構之間的合作關係,呈現為圖形。

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).


This investigation was prompted by interest in whether shifts in productivity based on geography are observed in the bibliometrics, informetrics and scientometrics areas.

One of the authors conducted a pilot study to determine whether there have been clear declines in North American contributions to the metrics literature base (Wolfram, 2008). The author found that there was indeed a notable relative decline in North American contributions and a sharp increase in European contributions.

Hood and Wilson (2001) examined the growth of literature of the metrics area. They provided an historical treatment of the development of these areas that included earlier studies of the field. In their research, literature associated with bibliometrics, informetrics and scientometrics was compared for the period 1968–2000. The authors noted that bibliometrics was still the most widely used term for metrics research.

More recently, Stock and Weber(2006) conducted a Web of Science search for records specifically including metrics terms and allied areas. They observed contributions had grown substantially since 1980.

Search parameters included the Boolean ORed result of bibliometrics, scientometrics, informetrics, cybermetrics and webometrics, in truncated form (e.g., webometri*), along with the phrases “citation analysis”, “link analysis” and “citation indexes”. ... These search results were ORed with the two primary journals that publish metrics research that are indexed by WoS, namely Scientometrics and the Journal of Informetrics.

A pair-wise comparison of all collaborations at the national and institutional levels was then conducted from which a cooccurrence matrix could be compiled.

Multidimensional scaling (MDS) analysis was used to visualize the relationships among countries. Because the data represent a type of similarity measure represented as a symmetric matrix, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006).

The recently developed visualization tool VOSviewer (van Eck &Waltman, 2010) was also used to provide an alternate visualization of the relationship outcomes. Like MDS, VOSviewer (http://www.vosviewer.com/) relies on a distance-based approach to mapping informetric relationships. Instead of using more traditional similarity measures to produce a normalized outcome for co-occurrences as used in MDS, relationships are based on association strengths, so the algorithm is somewhat different than PROXSCAL and, therefore, can produce different outcomes. Details of the comparison of different measures can be found in van Eck and Waltman (2009).

The network visualization software Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/) was used as well. Unlike the distance-based mapping of PROXSCAL and VOSviewer, Pajek produces directed or undirected network maps, with the strength of the relationships represented by the thickness of connecting lines between vertices on the map. Distances are used more for clarification, but proximities do not necessarily indicate a stronger relationship.

The search parameters retrieved 4404 publications.

Europe shows the highest levels of contribution, both in absolute and relative terms over the time period of the study. Growth patterns in absolute terms are nonlinear based on trend line analysis in MS Excel; however, the R-squared goodness-of-fit values for even the best fitting models (higher order polynomials) were never more than 0.95, indicating a less than desirable fit.

Relative contributions based on geographic divisions have been largely stable. An exception is Asia, which had an increasing relative contribution over the 22-year time frame of the study. Although North American contributions have continued to increase in absolute numbers, the relative contribution shows a slow average decline over time.

The top five journals listed for each continent demonstrated a regional preference for publication outlets from that region. So, for example, four of the top five journals for European publications were published in Europe, and four of the top five journal outlets for South America were South American. The exception to this was Asia. Four of the top five journals for Asian publications were European and one was North American. This outcome may be a reflection of the data extraction method, the indexing practices of WoS, or a preference during the study time frame for Asian scholars to publish in Western journals.

Seventy-five countries were represented in the record set.

The number of metrics papers published annually that represent collaborations between two or more countries has increased greatly since the mid-1990s. Prior to this time, the number of internationally collaborative papers ranged from 1 to 19 papers annually. Over the last decade this number has increased to a high of 96 papers in 2008.

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).

Sixteen of the institutions on the list are European, eight are North American, and one is Asian. The United States has the largest number of institutions represented (five), followed by Belgium (four – note: one institution merged with another institution to form a new entity).

There has been steady growth in inter-institutional collaboration over the 22 years. The mean number of collaborative institutional partners within the dataset has steadily increased from a low mean of 1.1 institutions per publication in 1987 to a high of 1.96 institutions per publication in 2007.

Europe, and in particular Western Europe, clearly dominates in the production of metrics literature. The United States continues to be the largest singular contributor, but this appears to be changing. North American contributions as a whole continue to increase, but represent a smaller percentage of worldwide production. European contributions have grown tremendously, especially during the last 5 years of the study period. This same period is marked by impressive growth from Asia.

It should be noted that WoS increased its coverage in 2008 by including more regional journals. These inclusions possibly could contribute to the increase in Asian contributions, but the observed growth for Asia was already evident prior to any such additions.

International and inter-institutional collaborations do not necessarily reveal strong geographic affinities, although the multiple institutional affiliations by a number of scholars associated with Flemish institutions do contribute to the strengthening of regional ties. Undoubtedly, the growth of the Internet and increasing availability of other telecommunication technologies have made these collaborations less distance dependent.

2014年1月23日 星期四

Egghe, L. (2012). Five years “Journal of Informetrics”. Journal of Informetrics, 6(3), 422-426.

Egghe, L. (2012). Five years “Journal of Informetrics”. Journal of Informetrics, 6(3), 422-426.

Journal of Informetrics (JOI)上發表的論文經過審查而具有良好模型與資料集合,主旨與資訊科學的基本量化方面有關的論文,範疇包含書目計量(bibliometrics)、科學計量(scientometrics)、網路計量(webometrics、cybermetrics)等廣義的資訊計量研究。最初的5卷裡,包含"letters to the editor",共計239篇論文,544位作者,每篇論文的平均作者數為2.276,其分布的主題如下表:
論文最多的兩個主題,引用分析(Citation analysis)和h指標(h-Type indices),便佔有一半以上的論文。

JOI最相近的期刊是Scientometrics,兩者都有相當多關於引用分析的論文。不同在於JOI有較多關於模型-理論(model-theoretic)以及網路議題(networking issue)方面的論文;Scientometrics則包含許多案例研究(case studies)的論文。

以一份相當年輕的期刊來說,JOI具有相當的影響力。利用Journal Citation Reports (JCR)提供的影響係數(impact factor)來觀察,JOI在2009與2010年都有很高的影響係數,在主題是圖書資訊學(Information and Library Science)的所有期刊中列於前三、四名。

將JOI最常引用的30種期刊以及最常被引用的30種期刊彼此間的引用關係,利用Gephi網絡分析軟體的ForceAtlas 2布局演算法繪製成網絡圖,使得彼此有引用關係的期刊在圖形上形成叢集。在JOI上方的期刊叢集是它經常引用的多元領域(multidisciplinary)期刊,右方的叢集是物理學相關期刊,下方是數學和資訊科學相關期刊,左方的期刊則以Research Policy和Research Evaluation為主。


JOI publishes refereed articles on fundamental quantitative aspects of information science. Accepted articles should contain good models and/or fundamental data sets.

The Journal covers the broad field of informetrics, including the field bibliometrics, scientometrics, webometrics and cybermetrics.

Specific topics can be described (non-exhaustively) as follows: informetric laws, modelling generalised bibliographies, aspects of inequality or concentration and diffusion, citation theory, linking theory (in general: social networks, including the Internet, citation and collaboration networks), downloads, indicators, evaluation techniques for scientific output (literature, scientists), evaluation techniques for documentary systems (information retrieval) including ranking theory, digital and classical library management, visualisation and mapping of science (individuals, fields, institutes, topics).

Over the five volumes there are 239 published articles with an average number of authors per article equal to 2.276. Here all articles and “letters to the editor” are taken into account. In total there are 544 (co-)authors.

The topics of these papers are described in Table 3.


We linked only one (main) topic to each paper. So a paper does not fall in two or more categories. Of course, sometimes, a paper could be linked to more than one category (e.g. papers on h-type indices can also deal with citation analysis) but we feel that Table 3 gives a rather accurate topical view.

It is clear that a bit more than 50% of the papers deal with citation analysis and/or h-type indices.

Finally the journal Scientometrics is closest to JOI in that it also publishes mainly on citation analysis. A difference between Scientometrics and JOI is that JOI publishes more model-theoretic papers and papers on networking issues while Scientometrics publishes more case studies.

These are very high numbers, certainly for a young journal. They are the highest for any “metrics” journal in the Journal Citation Reports (JCR) Subject Category Listing “Information and Library Science” (LIS).

The slight decrease of the IF from 2009 to 2010 is also noted by other LIS metrics journals, probably due to the fact that papers on the h-index (and related indices) have reached their maximum in terms of citations.

In fact JOI increased its relative impact from 2009 to 2010 since the value IF = 3.379 ranked JOI fourth out of 66 journals in the LIS Subject Category Listing while the value IF = 3.119 ranked JOI third out of 76 journals in the LIS Subject Category Listing.




The map is produced in Gephi, using the ForceAtlas 2 layout algorithm (http://webatlas.fr/tempshare/ForceAtlas2 Paper.pdf). It positions all journals with respect to the journals that they cite. Journals are mapped out using all of the citation links between them: related journals cluster together as they cite one another more frequently. Journals are selected for the map by virtue of being among the top 30 journals that JOI cites or the top 30 journals citing JOI in the mentioned period.

In this map, JOI sits in the centre with a core of informetrics journals, with branches leading to different clusters of research.

At the top are the large multidisciplinary journals (mainly cited by JOI rather than citing it); at the right, a group of physics journals with “Physics World” acting as a bridge from the informetrics journals. At the bottom of the map are mathematical and information science journals and at the left are “Research Policy” and “Research Evaluation”.

2014年1月18日 星期六

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

本研究利用社會網絡分析、共現分析(co-occurrence analysis)、叢集分析和詞語的頻率分析等多種分析技術,從Scientometrics期刊1978到2004年發表的1927筆論文資料,探討科學家合作網絡的結構特性、整個網絡上的合作領域以及個別的合作網絡、合作網絡上的合作中心(collaborative  center)。

過去的研究裡,Schubert (2002) 和 Dutt, Garg, & Bali (2003)都是針對國家間合作的巨觀層次。Kretschmer (2004) 認為巨觀和中觀(meso)層次的分析無法足夠地反映個人之間的合作趨勢,因此呼籲應在微觀層次的分析投注更多努力。

1927筆論文資料裡,單一作者的論文共有1052筆,所以仍稍占多數。作者數大於3的論文僅占非單一作者論文的13.71% (120/875),顯然研究Scientometrics的團隊規模都不大。發表3篇論文以及以上的高生產作者共計234人,其中有69.66%的作者曾發表與其他作者合作的論文。將這些作者間的合作關係表現成網絡,並利用Bibexcel對這個網絡上的節點進行叢集分析,共發現22個叢集。前兩個較大的叢集分別有15與14個科學家。網絡上最大的相連成分上共有15個叢集,共有合作經驗的高生產作者中的96位,占58.90%。合作網絡共有401條連結線,網絡密度為0.03,顯示Scientometrics領域的合作很鬆散。

對每一個節點計算它們的三種中心性,結果發現中心性和對應作者的生產力之間有很顯著的正相關,表示高生產力的作者同時也活躍在Scientometrics領域的合作網絡上。其中Glänzel的程度中心性最高,總共和其他18位作者有合作關係。

以詞語的頻率分析每個叢集的主題,最大的兩個叢集有類似的主題,但使用的研究方法略有不同。此外,研究主題為科學合作的四個叢集間幾乎沒有連結,同樣的情形也發生在研究科學與技術之間關係的四個叢集。

The structure of scientific collaboration networks in scientometrics is investigated at the level of individuals by using bibliographic data of all papers published in the international journal Scientometrics retrieved from the Science Citation Index (SCI) of the years 1978–2004.

Combined analysis of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of words is explored to reveal: (1) The microstructure of the collaboration network on scientists’ aspects of scientometrics; (2) The major collaborative fields of the whole network and of different collaborative sub-networks; (3) The collaborative center of the collaboration network in scientometrics.

Schubert [8] and Dutt etc. [9] presented international collaboration characteristics in the scientometrics community itself, focusing on country aspects at macro level.

Kretschmer [6] appealed to devote more efforts to investigations at micro level in the future because the knowledge at meso and macro level does not yet adequately reflect the trends in cooperation between individuals.

The study is based on bibliographic data retrieved from the Web of Science. The data contains all types of documents published in Scientometrics during 1978 to 2004.

In this study we have adapted an integrated procedure of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of title words.

Bibexcel is designed as a tool for manipulating bibliographic data, which is a free online-software published by Persson. In the present study, Bibexcel is used to do cooccurrence analysis and cluster analysis.

Following the methods of Otte & Rousseau [11], White [13] and Kretschmer & Aguillo [12], SNA was applied to display the microstructure of collaboration networks in scientometrics with Pajek.

Moreover, we used frequency analysis of title words to display the main collaborative field of different sub-networks. The software for frequency analysis is demo version of Wordsmith Tools published by Oxford University Press and available online.

There were 1927 documents published in Scientometrics during 1978 to 2004 (see Table 1).



From Table 1, we found that the pattern of co-authorship was still dominated by single-authored papers as the conclusion drawn by Dutt etc. [9].

While the number of multi-authored papers (the number of co-authors is more than 3) accounts for 13.71% only, which indicates that team size in scientometrics is not large.

In order to show the main structure of the network, each author must published 3 papers or more to be included in this integrated analysis. This threshold resulted in a total of 234 prolific authors publishing 3 or more papers during 1978 to 2004, among them there are 163 authors published co-authorship papers, accounting for 69.66% of the prolific authors.



Based on cluster analysis embedded in Bibexcel, we gained 22 clusters circled by solid lines (see Figure 1). We identified these clusters as sub-networks in the field of scientometrics.

The largest subnetwork is number 1 that has 15 collaborators, and the second largest one is number 2, which has 14 collaborators, and so on.

We noticed that there was totally 15 subnetworks connected with each other composing the largest central component, which had 96 numbers accounting for 58.90% of the prolific authors published co-authorship papers.

Density is an indicator for the general level of connectedness of the graph. ... In the present study, there are totally 401 links in the network, so the density of the network is 0.03, which indicates that the collaborative network in the field of scientometrics is very loose.

So an author who has high degree centrality must has collaborated with many other authors, which means the author is a central collaborator of the whole network. In the present study, Glänzel who has 18 co-workers is the central author of the whole network.

We found a positive and significant correlation between output of authors and the centrality measures (r=0.648, 0.437, 0.338 respectively at the 0.01 level, see Table 4) after investigating the correlations between output and the three centralities of the 125 authors in the 22 sub-networks, which indicated that most of the prolific authors are also active in collaboration network in the field of scientometrics.

We have also presented the main collaborative field of different sub-networks in scientometrics and found that the two biggest sub-networks have the similar collaborative topic with slightly methodological difference. In addition, we found an interesting phenomenon that four sub-networks dealing with scientific collaboration didn't collaborate with each other except sub-network 3 and 12. Moreover, four subnetworks studying technology and science never collaborated with each other at all.

2013年4月29日 星期一

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

information visualization

詞語地圖(term map)能夠將科學領域的結構視覺化。在這裡,詞語指的是能夠代表領域特定概念的詞(words)或片語(phrase),詞語地圖便是為了呈現出領域內重要的詞語之間的關係所產生的圖形。為了製作詞語地圖,本研究提出自動詞語確認(automatic term identification)的方法,以減少專家勞力並避免主觀判斷帶來的問題。考慮到從語料庫確認的詞語必須同時具有單元完整性(unithood)和主題相關性(termhood)兩方面的特質 (Kageura and Umino, 1996),本研究建議的方法包括三個階段:第一階段利用詞類標示器(part-of-speech tagger)產生出來的結果(Schmid,  1994; Schmid, 1995),抽取輸入語料內的名詞片語,做為候選詞語。第二階段比較候選詞語的出現頻率和候選詞語內的第一個詞與其餘部分的出現頻率,計算概似比(likelihood ratio) (Dunning, 1993),評估它們為完整語意單位(semantic unit)的程度,挑選單元完整性比較高的候選詞語。第三階段計算詞語的主題相關性是本研究的重要貢獻。在確認單元完整並與主題相關的詞語後,以每一對詞語之間相關強度(association strength) (Van Eck and Waltman 2009)的值代表它們之間的關係,利用VOS技術 (Van Eck and Waltman 2007a)產生詞語地圖。
本研究建議利用詞語在各主題上的分布傾向來估計每一個詞語的主題相關性。在本研究裡具有較高主題相關性的詞語是只與某一個或較少數主題有較強的關連的詞語。所以對於每一個詞語,本研究建議比較此一詞語在各主題上的分布情形與原先各主題的分布情形,如果差異較大便表示該詞語的主題相關性較高,也就是具有較高主題相關性的詞語對於少數的主題具有區辨力(discriminatory)。但由於每一篇文件都可能包括多個主題,無法單純地統計詞語在各主題上的分布情形以及各主題的分布情形,因此本研究利用機率式隱含語意分析(probabilistic latent semantic analysis, PLSA)的方式(Hofmann, 2001)估計各種詞語在各主題上的分布情形,並與原先各主題的分布情形相比較,找出分布偏向於少數主題的詞語。
評估自動化詞語確認的結果相當困難(Pazienza et al., 2005)。本研究為了評估詞語確認的結果,以15種 ISI主題分類為作業研究(operational research)的期刊,建立了該領域的詞語地圖並且以兩種方式進行評估:第一種方式是比較這種方法與沒有使用PLSA的詞語確認和利用詞語的出現頻率(frequency of occurrence)選取詞語等其他兩種方法的回收率(recall)與精確率(precision)。第二種方式則是由作業研究領域的專家對產生出來的詞語地圖進行品質審核。第一種評估方法的結果顯示除了在最高和最低的回收率以外,本研究建議的方法都比其他兩種方法能夠得到更高的精確率。在專家審查的結果則發現本研究產生的詞語地圖能夠表現出作業研究領域可分為以方法論為導向(methodology-oriented)及以應用為導向(application-oriented)的兩類研究主題,這個結果相當符合專家的想法。但是目前的結果也呈現出這個方法獲得意義較為廣泛的詞語、圖形上沒有包括某些主題以及有些主題非常相近的詞語在圖形上彼此間並不靠近等問題。

A term map is a map that visualizes the structure of a scientific field by showing the relations between important terms in the field.

To evaluate the proposed methodology, we use it to construct a term map of the field of operations research. The quality of the map is assessed by a number of operations research experts.

Other maps show relations between words or keywords based on co-occurrence data (e.g., Rip and Courtial 1984; Peters and Van Raan 1993; Kopcsa and Schiebel 1998; Noyons 1999; Ding et al. 2001). The latter maps are usually referred to as co-word maps.

By a term we mean a word or a phrase that refers to a domain-specific concept. Term maps are similar to co-word maps except that they may contain any type of term instead of only single-word terms or only keywords.

Selection of terms based on their frequency of occurrence in a corpus of documents typically yields many words and phrases with little or no domain-specific meaning. Inclusion of such words and phrases in a term map is highly undesirable for two reasons. First, these words and phrases divert attention from what is really important in the map. Second and even more problematic, these words and phrases may distort the entire structure shown in the map.

However, manual term selection has serious disadvantages as well. The most important disadvantage is that it involves a lot of subjectivity, which may introduce significant biases in a term map. Another disadvantage is that it can be very labor-intensive.

Given a corpus of documents, we first identify the main topics in the corpus. This is done using a technique called probabilistic latent semantic analysis (Hofmann 2001). Given the main topics, we then identify in the corpus the words and phrases that are strongly associated with only one or only a few topics. These words and phrases are selected as the terms to be included in a term map.

An important property of the proposed methodology is that it identifies terms that are not only domain-specific but that also have a high discriminatory power within the domain of interest. This is important because terms with a high discriminatory power are essential for visualizing the structure of a scientific field.

We define unithood as the degree to which a phrase constitutes a semantic unit. Our idea of a semantic unit is similar to that of a collocation (Manning and Schu¨tze 1999). Hence, a semantic unit is a phrase consisting of words that are conventionally used together. The meaning of the phrase typically cannot be fully predicted from the meaning of the individual words within the phrase.

We define termhood as the degree to which a semantic unit represents a domain-specific concept.

Linguistic approaches are mainly used to identify phrases that, based on their syntactic form, can serve as candidate terms.

Statistical approaches are used to measure the unithood and termhood of phrases.

Most terms have the syntactic form of a noun phrase (Justeson and Katz 1995; Kageura and Umino 1996). Linguistic approaches to automatic term identification typically rely on this property. These approaches identify candidate terms using a linguistic filter that checks whether a sequence of words conforms to some syntactic pattern. Different researchers use different syntactic patterns for their linguistic filters (e.g., Bourigault 1992; Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995; Frantzi et al. 2000).

Statistical approaches to measure unithood are discussed extensively by Manning and Schu¨tze (1999). The simplest approach uses frequency of occurrence as a measure of unithood (e.g., Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995). More advanced approaches use measures based on, for example, (pointwise) mutual information (e.g., Church and Hanks 1990; Damerau 1993; Daille et al. 1994) or a likelihood ratio (e.g., Dunning 1993; Daille et al. 1994). Another statistical approach to measure unithood is the C-value (Frantzi et al. 2000). The NC-value (Frantzi et al. 2000) and the SNC-value (Maynard and Ananiadou 2000) are extensions of the C-value that measure not only unithood but also termhood. Other statistical approaches to measure termhood can be found in the work of, for example, Drouin (2003) and Matsuo and Ishizuka (2004). In the field of machine learning, an interesting statistical approach to measure both unithood and termhood is proposed by Wang et al. (2007).

Termhood is measured as the degree to which the occurrences of a semantic unit are biased towards one or more topics.

In the first step of our methodology, we use a linguistic filter to identify noun phrases. We first assign to each word occurrence in the corpus a part-of-speech tag, such as noun, verb, or adjective. The appropriate part-of-speech tag for a word occurrence is determined using a part-of-speech tagger developed by Schmid (1994, 1995). We use this tagger because it has a good performance and because it is freely available for research purposes.

The most common approach to measure unithood is to determine whether a phrase occurs more frequently than would be expected based on the frequency of occurrence of the individual words within the phrase.

To measure the unithood of a noun phrase, we first count the number of occurrences of the phrase, the number of occurrences of the phrase without the first word, and the number of occurrences of the first word of the phrase. In a similar way as Dunning (1993), we then use a so-called likelihood ratio to compare the first number with the last two numbers.

The main idea of the third step of our methodology is to measure the termhood of a semantic unit as the degree to which the occurrences of the unit are biased towards one or more topics.

To measure the degree to which the occurrences of semantic unit uk, where k (belongs to) {1,…,K}, are biased towards one or more topics, we use two probability distributions, namely the distribution of semantic unit uk over the set of all topics and the distribution of all semantic units together over the set of all topics. These distributions are denoted by, respectively, P(tj | uk) and P(tj), where j (belongs to) {1,…, J}. ... The dissimilarity between the two distributions indicates the degree to which the occurrences of uk are biased towards one or more topics. We use the dissimilarity between the two distributions to measure the termhood of uk.

For example, if the two distributions are identical, the occurrences of uk are unbiased and uk most probably does not represent a domain-specific concept. If, on the other hand, the two distributions are very dissimilar, the occurrences of uk are strongly biased and uk is very likely to represent a domain-specific concept.

The dissimilarity between two probability distributions can be measured in many different ways. One may use, for example, the Kullback–Leibler divergence, the Jensen–Shannon divergence, or a chi-square value.

In (3), termhood (uk) is calculated as the negative entropy of this distribution. Notice that termhood (uk) is maximal if P(tj | uk) = 1 for some j and that it is minimal if P(tj | uk) = P(tj) for all j. In other words, termhood (uk) is maximal if the occurrences of uk are completely biased towards a single topic, and termhood (uk) is minimal if the occurrences of uk do not have a bias towards any topic.

In order to allow for a many-to-many relationship between corpus segments and topics, we make use of probabilistic latent semantic analysis (PLSA) (Hofmann 2001).

It was originally introduced as a probabilistic model that relates occurrences of words in documents to so-called latent classes. In the present context, we are dealing with semantic units and corpus segments instead of words and documents, and we interpret the latent classes as topics.

PLSA assumes that each occurrence of a semantic unit in a corpus segment is independently generated according to the following probabilistic process. First, a topic t is drawn from a probability distribution P(tj), where j (belongs to) {1,…,J}. Next, given t, a corpus segment s and a semantic unit u are independently drawn from, respectively, the conditional probability distributions P(si | t), where i (belongs to) {1,…,I}, and P(uk | t), where k (belongs to) {1,…,K}. This then results in the occurrence of u in s.

P(si, uk) =  sum (from j=1 to J) P(tj)P(si|tj)P(uk|tj)

We estimate these parameters using data from the corpus. Estimation is based on the criterion of maximum likelihood. The log-likelihood function to be maximized is given by
L = sum(from i=1 to I) sum(from k=1 to K) nik log P(si, uk)
We use the EM algorithm discussed by Hofmann (1999, Sect. 3.2) to perform the maximization of this function.

After estimating the parameters of PLSA, we apply Bayes’ theorem to obtain a probability distribution over the topics conditional on a semantic unit. This distribution is given by

P(tj|uk) = P(tj)P(uk|tj) / sum (from j=1 to J) (P(P(tj)P(uk|tj))

In a similar way as discussed earlier, we use the dissimilarity between the distributions P(tj | uk) and P(tj) to measure the termhood of uk.

We first selected a number of OR journals. This was done based on the subject categories of Thomson Reuters. The OR field is covered by the category Operations Research & Management Science. Since we wanted to focus on the core of the field, we selected only a subset of the journals in this category. More specifically, a journal was selected if it belongs to the category Operations Research & Management Science and possibly also to the closely related category Management and if it does not belong to any other category. This yielded 15 journals, which are listed in the first column of Table 1.

In the first step of our methodology, the linguistic filter identified 2662 different noun phrases. In the second step, the unithood of these noun phrases was measured. 203 noun phrases turned out to have a rather low unithood and therefore could not be regarded as semantic units. ... The other 2459 noun phrases had a sufficiently high unithood to be regarded as semantic units.

In the third and final step of our methodology, the termhood of these semantic units was measured. To do so, each title-abstract pair in the corpus was treated as a separate corpus segment. For each combination of a semantic unit uk and a corpus segment si, it was determined whether uk occurs in si (nik = 1) or not (nik = 0). Topics were identified using PLSA. This required the choice of the number of topics J. Results for various numbers of topics were examined and compared. Based on our own knowledge of the OR field, we decided to work with J = 10 topics.

The evaluation of a methodology for automatic term identification is a difficult issue. There is no generally accepted standard for how evaluation should be done. We refer to Pazienza et al. (2005) for a discussion of the various problems.

We first perform an evaluation based on the well-known notions of precision and recall. We then perform a second evaluation by constructing a term map and asking experts to assess the quality of this map.

Precision is the number of correctly identified terms divided by the total number of identified terms.

Recall is the number of correctly identified terms divided by the total number of correct terms.

Unfortunately, because the total number of correct terms in the OR field is unknown, we could not calculate the true recall. This is a well-known problem in the context of automatic term identification (Pazienza et al. 2005).

To circumvent this problem, we defined recall in a slightly different way, namely as the number of correctly identified terms divided by the total number of correct terms within the set of all semantic units identified in the second step of our methodology. Recall calculated according to this definition provides an upper bound on the true recall. However, even using this definition of recall, the calculation of precision and recall remained problematic. The problem was that it is very time-consuming to manually determine which of the 2459 semantic units identified in the second step of our methodology are correct terms and which are not. We solved this problem by estimating precision and recall based on a random sample of 250 semantic units.

It is clear from the figure that our methodology outperforms the two simple alternatives. Except for very low and very high levels of recall, our methodology always has a considerably higher precision than the variant of our methodology that does not make use of PLSA.

A term map is a map, usually in two dimensions, that shows the relations between important terms in a scientific field. Terms are located in a term map in such a way that the proximity of two terms reflects their relatedness as closely as possible. That is, the smaller the distance between two terms, the stronger their relation. The aim of a term map usually is to visualize the structure of a scientific field.

It turned out that, out of the 2459 semantic units identified in the second step of our methodology, 831 had the highest possible termhood value. This means that, according to our methodology, 831 semantic units are associated exclusively with a single topic within the OR field. We decided to select these 831 semantic units as the terms to be included in the term map. This yielded a coverage of 97.0%, which means that 97.0% of the title-abstract pairs in the corpus contain at least one of the 831 terms to be included in the term map.

The term map of the OR field was constructed using a procedure similar to the one used in our earlier work (Van Eck and Waltman 2007b). This procedure relies on the association strength measure (Van Eck and Waltman 2009) to determine the relatedness of two terms, and it uses the VOS technique (Van Eck and Waltman 2007a) to determine the locations of terms in the map.

The most serious criticism on the results of the automatic term identification concerned the presence of a number of rather general terms in the map.

Another point of criticism concerned the underrepresentation of certain topics in the term map. There were three experts who raised this issue. One expert felt that the topic of supply chain management is underrepresented in the map. Another expert stated that he had expected the topic of transportation to be more visible. The third expert believed that the topics of combinatorial optimization, revenue management, and transportation are underrepresented.

As discussed earlier, when we were putting together the corpus, we wanted to focus on the core of the OR field and we therefore only included documents from a relatively small number of journals. This may for example explain why the topic of transportation is not clearly visible in the map.

When asked to divide the OR field into a number of smaller subfields, most experts indicated that there are two natural ways to make such a division. On the one hand, a division can be made based on the methodology that is being used, such as decision theory, game theory, mathematical programming, or stochastic modeling. On the other hand, a division can be made based on the area of application, such as inventory control, production planning, supply chain management, or transportation. There were two experts who noted that the term map seems to mix up both divisions of the OR field. According to these experts, one part of the map is based on the methodology-oriented division of the field, while the other part is based on the application-oriented division.

The experts pointed out that sometimes closely related terms are not located very close to each other in the map. One of the experts gave the terms inventory and inventory cost as an example of this problem. In many cases, a problem such as this is probably caused by the limited size of the corpus that was used to construct the map. In other cases, the problem may be due to the inherent limitations of a two-dimensional representation.

Our main contribution consists of a methodology for automatic identification of terms in a corpus of documents. Using this methodology, the process of selecting the terms to be included in a term map can be automated for a large part, thereby making the process less labor-intensive and less dependent on expert judgment. Because less expert judgment is required, the process of term selection also involves less subjectivity.

In general, we are quite satisfied with the results that we have obtained. The precision/recall results clearly indicate that our methodology outperformed two simple alternatives. In addition, the quality of the term map of the OR field constructed using our methodology was assessed quite positively by five experts in the field. However, the term map also revealed a shortcoming of our methodology, namely the incorrect identification of a number of general noun phrases as terms.

As scientific fields tend to overlap more and more and disciplinary boundaries become more and more blurred, finding an expert who has a good overview of an entire domain becomes more and more difficult. This poses serious difficulties for any bibliometric method that relies on expert knowledge.

2013年4月14日 星期日

Waltman, L., van Eck, N. J., & Noyons, E. (2010). A unified approach to mapping and clustering of bibliometric networks. Journal of Informetrics, 4(4), 629-635.

Waltman, L., van Eck, N. J., & Noyons, E. (2010). A unified approach to mapping and clustering of bibliometric networks. Journal of Informetrics, 4(4), 629-635.

information visualization

本研究提出整合書目計量網絡(bibliometric networks)的映射(mapping)與叢集(clustering)的技術。映射與叢集技術經常一起用於分析書目網絡的結構,以了解科學領域的重要研究主題、研究主題之間的關係和領域的發展等問題。在過去對於書目網絡的相關研究中,有些研究建構一個映射圖來呈現網絡上的節點,並且在圖上呈現節點的叢集情形,例如McCain (1990)、White & Griffith (1981)、Leydesdorff & Rafols (2009)和Van Eck, Waltman, Dekker, & Van den Berg (in press);有些研究則先對節點進行叢集,然後建構一個映射圖來呈現節點的叢集,例如Small, Sweeney, & Greenlee (1985)和Noyons, Moed, & Van Raan, (1999);第三種方式則是先建構一個呈現節點的映射圖,再利用節點在映射圖上的座標進行叢集,例如Boyack, Klavans, & Börner (2005)和Klavans & Boyack (2006)。在書目計量學和科學計量學的研究裡常使用的映射與叢集技術組合是以多維縮放(multidimensional scaling)和 階層叢集(hierarchical clustering)技術的組合,著名的早期研究有McCain (1990)、Peters &Van Raan (1993)、Small et al., (1985)和White & Griffith (1981)。其他知名的映射技術還有經常配合尋徑者網路縮放方法(pathfinder network scaling)的Kamada and Kawai (1989)映射演算法,例如 Chen (1999)、de Moya-Anegón et al. (2007)和White (2003)等,Boyack等研究者提出的VxOrd(Boyack et al., 2005; Klavans & Boyack, 2006)和Van Eck等研究者提出的VOS (Van Eck et al., in press)也都是常被使用的映射技術。在叢集方面,除了階層式叢集以外,因素分析(factor analysis)也常被使用,例如de Moya-Anegón et al. (2007)、Leydesdorff & Rafols (2009)和Zhao & Strotmann (2008)等研究,近年來書目計量學和科學計量學的研究裡經常被應用的技術是建立在Newman and Girvan (2004)提出的模組性函數(modularity function)的叢集技術,例如Chen & Redner (2010)、Lambiotte & Panzarasa (2009)、Schubert & Soós (2010)、Takeda & Kajikawa (2009)、Wallace, Gingras, & Duhon, (2009)和Zhang, Liu, Janssens, Liang, & Glänzel (2010)。然而正如以上的分析,映射與叢集技術一起用於分析書目網絡的結構的技術,雖然極為相關,但是大多是獨立發展。本研究便是基於這個問題,藉由對於過去發展的VOS映射技術以及以模組性為基礎的叢集技術引導出一致的原則,建立這兩種技術的關連(relation),來進行整合。另一個整合映射和叢集技術的研究Noack (2009)則定義了一個參數化的目標函數(a parameterized objective function)來描述一類的映射技術, 並且證明以模組化為基礎的叢集技術也可以納入這個目標函數,因此可以建立映射和叢集技術之間的關係。本研究與Noack(2009)的不同在於本研究提出的方法直接建立VOS映射技術和模組性為基礎的叢集技術之間的關係,而不是透過目標函數做為映射和叢集技術之間的關係,並且也包含一個權重因素(weighing factor),最後本研究的方法利用解析度(resolution)參數來解決模組性為基礎的叢集技術在解析度上的問題。為了驗證這個技術的可行性,本研究並且以資訊科學在1999到2008年間最常被引用的1242筆文獻進行映射和叢集,利用書目耦合和共被引次數的總和來估計文獻間的關連程度,產生的結果圖形上可以觀察到在資訊科學的結構中包含資訊尋求和檢索(information seeking and retrieval)以及資訊計量學(informetrics)兩個大的次領域,這個結果與其他以資訊科學為分析對象的書目計量研究相似。

In bibliometric and scientometric research, a lot of attention is paid to the analysis of networks of, for example, documents, keywords, authors, or journals. Mapping and clustering techniques are frequently used to study such networks.The aim of these techniques is to provide insight into the structure of a network. The techniques are used to address questions such as:
• What are the main topics or the main research fields within a certain scientific domain?
• How do these topics or these fields relate to each other?
• How has a certain scientific domain developed over time?
To satisfactorily answer such questions, mapping and clustering techniques are often used in a combined fashion.

One approach is to construct a map in which the individual nodes in a network are shown and to display a clustering of the nodes on top of the map, for example by marking off areas in the map that correspond with clusters (e.g., McCain, 1990; White & Griffith, 1981) or by coloring nodes based on the cluster to which they belong (e.g., Leydesdorff & Rafols, 2009; Van Eck, Waltman, Dekker, & Van den Berg, in press).

Another approach is to first cluster the nodes in a network and to then construct a map in which clusters of nodes are shown. This approach is for example taken in the work of Small et al. (e.g., Small, Sweeney, & Greenlee, 1985) and in earlier work of our own institute (e.g., Noyons, Moed, & Van Raan, 1999).

A third approach is to first construct a map in which the individual nodes in a network are shown and to then cluster the nodes based on their coordinates in the map (e.g., Boyack, Klavans, & Börner, 2005; Klavans & Boyack, 2006).

In the bibliometric and scientometric literature, the most commonly used combination of a mapping and a clustering technique is the combination of multidimensional scaling and hierarchical clustering (for early examples, see McCain, 1990; Peters&Van Raan, 1993; Small et al., 1985; White&Griffith, 1981).

A popular alternative to multidimensional scaling is the mapping technique of Kamada and Kawai (1989); (see e.g. Leydesdorff & Rafols, 2009; Noyons & Calero-Medina, 2009), which is sometimes used together with the pathfinder network technique (Schvaneveldt, Dearholt, & Durso, 1988; see e.g. Chen, 1999; de Moya-Anegón et al., 2007; White, 2003). Two other alternatives to multidimensional scaling are the VxOrd mapping technique (e.g., Boyack et al., 2005; Klavans & Boyack, 2006) and our own VOS mapping technique (e.g., Van Eck et al., in press).

Factor analysis, which has been used in a large number of studies (e.g., de Moya-Anegón et al., 2007; Leydesdorff & Rafols, 2009; Zhao & Strotmann, 2008), may be seen as a kind of clustering technique and, consequently, as an alternative to hierarchical clustering. Another alternative to hierarchical clustering is clustering based on the modularity function of Newman and Girvan (2004); (see e.g. Wallace, Gingras, & Duhon, 2009; Zhang, Liu, Janssens, Liang, & Glänzel, 2010).

In bibliometric and scientometric research, modularity-based clustering has been used in a number of recent studies (Chen & Redner, 2010; Lambiotte & Panzarasa, 2009; Schubert & Soós, 2010; Takeda & Kajikawa, 2009; Wallace et al., 2009; Zhang et al., 2010).

As we have discussed, mapping and clustering techniques have a similar objective, namely to provide insight into the structure of a network, and the two types of techniques are often used together in bibliometric and scientometric analyses. However, despite their close relatedness, mapping and clustering techniques have typically been developed separately from each other.

In our view, when a mapping and a clustering technique are used together in the same analysis, it is generally desirable that the techniques are based on similar principles as much as possible. This enhances the transparency of the analysis and helps to avoid unnecessary technical complexity. Moreover, by using techniques that rely on similar principles, inconsistencies between the results produced by the techniques can be avoided.

In this paper, we propose a unified approach to mapping and clustering of bibliometric networks. We show how a mapping and a clustering technique can both be derived from the same underlying principle. In doing so, we establish a relation between on the one hand the VOS mapping technique (Van Eck &Waltman, 2007; Van Eck et al., in press) and on the other hand clustering based on a weighted and parameterized variant of the well-known modularity function of Newman and Girvan (2004).

It follows from (6) and (7) that our proposed clustering technique can be seen as a kind of weighted variant of modularity-based clustering (see Appendix B for a further discussion). However, unlike modularity-based clustering, our clustering technique has a resolution parameter . This parameter helps to deal with the resolution limit problem (Fortunato & Barthélemy, 2007) of modularity based clustering. Due to this problem, modularity-based clustering may fail to identify small clusters. Using our clustering technique, small clusters can always be identified by choosing a sufficiently large value for the resolution parameter .

The above result showing how mapping and clustering can be performed in a unified and consistent way resembles to some extent a result derived by Noack (2009). Noack defined a parameterized objective function for a class of mapping techniques (referred to as force-directed layout techniques by Noack). This class of mapping techniques includes for example the well-known technique of Fruchterman and Reingold (1991). Noack showed that his parameterized objective function subsumes the modularity function of Newman and Girvan (2004). In this way, Noack established a relation between on the one hand a class of mapping techniques and on the other hand modularity-based clustering.

First, the result of Noack does not directly relate well-known mapping techniques such as the one of Fruchterman and Reingold to modularity-based clustering. Instead, Noack’s result shows that the objective functions of some well-known mapping techniques and the modularity function of Newman and Girvan are special cases of the same parameterized function. Our result establishes a direct relation between a mapping technique that has been used in various applications, namely the VOS mapping technique, and a clustering technique.

Second, the mapping and clustering techniques considered by Noack and the ones that we consider differ from each other by a weighing factor. This is the weighing factor given by (7).

Third, the clustering technique considered by Noack is unparameterized, while our clustering technique has a resolution parameter.

In Fig. 1, we show a combined mapping and clustering of the 1242 most frequently cited publications that appeared in the field of information science in the period 1999–2008. The mapping and the clustering were produced using our unified approach.

For these publications, we determined the number of co-citation links and the number of bibliographic coupling links. These two types of links were added together and served as input for both our mapping technique and our clustering technique.

The combined mapping and clustering shown in Fig. 1 provides an overview of the structure of the field of information science. The left part of the map represents what is sometimes referred to as the information seeking and retrieval (ISR) subfield (Åström, 2007), and the right part of the map represents the informetrics subfield.

The clustering shown in Fig. 1 consists of 25 clusters. The distribution of the number of publications per cluster has a mean of 49.7 and a standard deviation of 31.5.

2013年3月27日 星期三

Cobo, M. J., López‐Herrera, A. G., Herrera‐Viedma, E., & Herrera, F. (2012). SciMAT: A new science mapping analysis software tool. Journal of the American Society for Information Science and Technology.

Cobo, M. J., López‐Herrera, A. G., Herrera‐Viedma, E., & Herrera, F. (2012). SciMAT: A new science mapping analysis software tool. Journal of the American Society for Information Science and Technology.

information visualization


本論文介紹科學映射分析(science mapping analysis)工具SciMAT的功能與應用。根據Börner et al., (2003)和Cobo et al., (2011b)等研究,科學映射分析的流程可以分成以下的步驟:1) 資料檢索 (data retrieval)、2) 資料前處理 (data preprocessing)、3) 網路資訊抽取 (network extraction)、4) 網路資訊正規化 (network normalization)、5) 映射 (mapping)、6) 分析 (analysis)以及7) 視覺化 (visualization)。特別要說明的是「資料前處理」是處理原始資料的重複和錯誤、區分時段(time slicing)以及網路資料縮減等工作,是決定科學映射分析能否得到良好結果的重要步驟之一。「網路資訊抽取」則是從論文的書目資料裡建立分析項目之間的關連,包括共現(co-occurrence)、耦合(coupling)和直接連結(direct linkage)等關係。兩個分析項目的共現關係取決於它們是否共同出現在一組文件內以及共同出現的次數;文件間的耦合關係則建立於它們是否具有共同的項目以及其數量大小,作者及期刊間的耦合關係則由屬於他們的文件的共同項目聚集而成;直接連結則是文件與它們的參考文獻之間的引用關係。運用不同的分析項目以及不同的關係可以對科學研究領域進行各種面向的分析,例如以文件中共同出現的作者所抽取的共同作者關係建立的網絡可以分析科學研究領域的社會結構(social structure);對於由詞語在文件內的共現關係所建構的詞語共現網絡進行分析則可以得知領域的概念結構(conceptual structure)和所處理的主要概念;經由文獻引用所產生的共被引關係和書目耦合關係則可以用來分析科學研究領域的知識結構(intellectual structure)。透過上面對於科學映射分析流程的分析,可以知道一個科學映射分析工具最好能夠具備以下的特性:a) 包含多種不同的模組來處理科學映射工作流程中的各個步驟;b) 具備強大的消除重複模組;c) 能夠建構各種書目計量的大型網絡;d) 具有良好的視覺化技術;e)輸出結果應該包含書目計量的測量結果與指標。本研究所提出的SciMAT工具具備上述的各種特性。SciMAT包含三個重要的模組:知識庫(knowledge base)、工作流程的配置以及測量結果與映射圖的視覺化模組。SciMAT的知識庫模組提供分析者匯入各種書目來源的檢索結果,將文件的作者、關鍵詞、期刊和參考文獻等各種資料儲存於知識庫內。運用此知識庫提供的功能,分析者能夠進行編輯與前處理等改善資料品質的工作以獲得更好的分析結果。SciMAT的工作流程配置模組循序漸進地設定分析的時間區段、分析的項目單位和關係、使用資料的次數閾值、進行資料正規化的相似性測量方式、叢集方式以及網絡分析、成效分析、時間分析和歷時性分析等相關參數。視覺化模組可以針對每個分析時段(period)提供詳細的網路圖、策略圖表以及相關的書目計量測量結果,也能夠提供代表研究主題(theme)的叢集在不同時段的演進情形等歷時性(longitudinal)的圖表

The general workflow in a science mapping analysis has different steps (Börner et al., 2003; Cobo et al., 2011b) (see Figure 1): data retrieval, data preprocessing, network extraction, network normalization, mapping, analysis, and visualization. At the end of this process, the analyst has to interpret and obtain conclusions from the results.

Usually, the data retrieved from the bibliographic sources contain errors, so a preprocessing process must be applied first. In fact, the preprocessing step is one of the most important to obtain good results in science mapping analysis. Different preprocessing processes can be applied to the raw data, such as detecting duplicate and misspelled items, time slicing, data reduction, and network reduction (for more information, see Cobo et al., 2011b).

A co-occurrence relation is established between two units (authors, terms, or references) when they appear together in a set of documents; that is, when they co-occur throughout the corpus.

A coupling relation is established between two documents when they have a set of units (authors, terms, or references) in common. Furthermore, the coupling can be established using a higher level unit of aggregation, such as authors or journals. That is, a coupling between two authors or journals can be established by counting the units shared by their documents (using the author’s or journal’s oeuvres).

Finally, a direct linkage establishes a relation between documents and references, particularly a citation relation.

In addition, different aspects of a research field can be analyzed depending on the units of analysis used and the kind of relation selected (Cobo et al., 2011b).

For example, using the authors, a coauthor or coauthorship analysis can be performed to study the social structure of a scientific field (Gänzel, 2001; Peters & van Raan, 1991).

Using terms or words, a co-word (Callon, Courtial, Turner, & Bauin, 1983) analysis can be performed to show the conceptual structure and the main concepts dealt with by a field.

Cocitation (Small, 1973) and bibliographic coupling (Kessler, 1963) are used to analyze the intellectual structure of a scientific research field.

We therefore think it would be desirable to develop a science mapping software tool that satisfies the following requirements: (a) it should incorporate modules to carry out all the steps of the science mapping workflow, (b) it should present a powerful de-duplicating module, (c) it should be able to build a large variety of bibliometric networks, (d) it should be designed with good visualization techniques, and (e) it should enrich the output with bibliometric measures.

SciMAT generates a knowledge base from a set of scientific documents where the relations of the different entities related to each document (authors, keywords, journal, references, etc.) are stored. This structure helps the analyst to edit and preprocess the knowledge base to improve the quality of the data and, consequently, obtain better results in the science mapping analysis.

Taking into account the GUI, there are three important modules: (a) a module dedicated to the management of the knowledge base and its entities, (b) a module (wizard) responsible for configuring the science mapping analysis, and (c) a module to visualize the generated results and maps. These modules allow the analyst to carry out the different steps of the science mapping workflow.

Regarding its functionalities, the module to manage the knowledge base is responsible for building the knowledge base, importing the raw data from different bibliographical sources, and cleaning and fixing the possible errors in the entities. It can be considered as a first stage in the preprocessing step.

As shown, the workflow is divided into four main stages: (a) to build the data set, (b) to create and normalize the network, (c) to apply a cluster algorithm to get the map, and (d) to perform a set of analyses. These stages and their respective steps are described below:
1. Build the data set: At this stage, the user can configure the periods of time used in the analysis (select the periods), the aspects that he or she wants to analyze (select the unit of analysis:  the conceptual (using terms or words), social (using authors), and intellectual (using references) aspects), and the portion of the data that has to be used (to filter the data using a minimum frequency as a threshold).
2. Create and normalize the network: At this stage, the network is built using co-occurrence or coupling relations or, indeed, aggregating coupling. Then, the network is filtered to keep only the most representative items. Finally, a normalization process is performed using a similarity measure (association strength (Coulter et al., 1998; van Eck &Waltman, 2007), Equivalence Index (Callon et al., 1991), Inclusion Index, Jaccard Index (Peters & van Raan, 1993), and Salton’s cosine (Salton & McGill, 1983).
3. Apply a clustering algorithm to get the map and its associated clusters or subnetworks: At this stage, the clustering algorithm used to build the map has to be selected. Different clustering methods are available in SciMAT, such as the Simple Centers Algorithm (Cobo et al., 2011a; Coulter et al., 1998), Single-linkage (Small & Sweeney, 1985), and variants such as Complete-linkage, Average-linkage, and Sum-linkage.
4. Apply a set of analyses: The final step of the wizard consists of selecting the analyses to be performed on the generated map.
(a) Network analysis: By default, SciMAT adds Callon’s density and centrality (Callon et al., 1991; Cobo et al., 2011a) as network measures to each detected cluster in each selected period. Callon’s centrality measures the degree of interaction of a network with other networks, and it can be understood as the external cohesion of the network. ... Callon’s density measures the internal strength of the network, and it can be understood as the internal cohesion of the network. ... These measures are useful to categorize the detected clusters of a given period in a strategic diagram (Cobo et al., 2011a).
(b) Performance analysis: SciMAT is able to assess the output according to several performance and quality measures. To do that, it incorporates into each cluster a set of documents using a document mapper function and then calculates the performance based on quantitative and qualitative measures (using citation-based measures, number of documents, etc.).
(c) Temporal analysis or longitudinal analysis: This allows the user to discover the conceptual, social, or intellectual evolution of the field. SciMAT is able to build an evolution map to detect the evolution areas (Cobo et al., 2011a) and an overlapping items graph (Price & Gürsey, 1975; Small, 1977) across the periods analyzed. Furthermore, SciMAT allows the user to choose different measures to calculate the weight of the “evolution nexus” (Cobo et al., 2011a) between the items of two consecutive periods, such as association strength (Coulter et al., 1998; van Eck & Waltman, 2007), Equivalence Index (Callon et al., 1991), Inclusion Index, Jaccard’s Index (Peters & van Raan, 1993), and Salton’s cosine (Salton & McGill, 1983).

At the end of all the steps in the wizard, the map would be built using the selected configuration. Then, the results would be saved to a file, and the visualization module loaded. The visualization module has two views: Longitudinal and Period.

The Period view (see Figure 12) shows detailed information for each period, its strategic diagram, and for each cluster, the bibliometric measures, the network, and their associated nodes.

Finally, in the Longitudinal view the overlapping map and evolution map are shown. This view helps us to detect the evolution of the clusters throughout the different periods, and study the transient and new items of each period and the items shared by two consecutive periods.

Taking into account quantitative measures such as the number of documents associated with each theme (cluster), we can discover where the fuzzy community has been employing a great effort (e.g., H-INFINITY-CONTROL, FUZZY-CONTROL, T-NORM, etc.). Similarly, taking into account the qualitative measure, we could identify the themes with a greater impact; that is, the themes that have been highly cited.

Combining the units of analysis and the bibliographic relations among them, SciMAT can extract 20 kinds of bibliographic networks, including the common bibliographic networks used in the literature, such as coauthor (Gänzel, 2001; Peters & van Raan, 1991), bibliographic coupling (Kessler, 1963), journal bibliographic coupling (Small & Koenig, 1977), author bibliographic coupling (Zhao & Strotmann, 2008), cocitation (Small, 1973), journal cocitation (McCain, 1991), author cocitation (White & Griffith, 1981), and co-word (Callon et al., 1983).

2013年3月19日 星期二

Klavans, R., & Boyack, K. W. (2006). Identifying a better measure of relatedness for mapping science. Journal of the American Society for Information Science and Technology, 57(2), 251-263.

Klavans, R., & Boyack, K. W. (2006). Identifying a better measure of relatedness for mapping science. Journal of the American Society for Information Science and Technology57(2), 251-263.

information visualization


本研究提出相關性(relatedness)測量的評量架構,並利用這個評量架構判定六種交互引用(intercitation)和四種共被引(cocitation)的相關性測量方式以及應用到視覺化演算法的結果,六種應用於交互引用的相關性測量方式包括原始的引用次數、cosine指標、Jaccard指標、Pearson相關係數、Pudovkin & Fuseler (1995)和Pudovkin & Garfield (2002)根據期刊引用應用所提出的相關因素(relatedness factor)、以及本研究提出的由cosine指標減去期望的cosine值(expected cosine value)的K50指標,四種應用於交互引用的相關性測量方式則有原始的共被引次數、cosine指標、Pearson相關係數和K50指標。本研究所提出來的評量架構是以一組已經分類的物件為基礎,評估各種相關性測量方式的準確度(accuracy)、覆蓋率(coverage)、可擴展性(scalablity)和強健性(robustness)。例如對期刊間的相關性進行測量時,可以利用ISI的期刊分類為評估物件間相關性的基礎。準確度是指能夠正確地判斷對象間是否相關,可以再區分成區域準確度(local accuracy)與整體準確度(global accuracy),區域準確度是指物件與其他最接近物件是否能夠被正確地放置與排序的趨勢,也就是在同一分類的物件是否具有比在不同分類的物件更高的相關性,整體準確度是指分類之間的位置與排序等關係。覆蓋率則是指的是某一個閾值(threshold)以上的相關性所得到的正確分類結果占所有應有的分類結果的比例。可擴展性是指這種測量方式能否應用於非常大型的資料集合,與測量方式的計算量有關。強健性則是指將相關性測量的結果應用到視覺化演算法進行維度縮減(dimensional reduction)處理後,物件在產生圖形上的映射點間的相關性能否保留原先測量方式的相關性的關係。這四種評估指標彼此間有所關連,例如較大的覆蓋率通常會得到較不準確的結果;而如果希望得到較準確的結果,採用較多計算量的測量方式便無法達到較好的可擴展性;而在維度縮減後,也可能導致準確度變差;最後,以交互引用資料做為相關性測量方式的輸入,能夠利用最近期的資料,獲得較準確的結果,但以共被引資料做為輸入,則能夠包含不在分析期刊中的來源。本研究以2000年ISI的SCIE(science citation index extended)和SSCI(social science citation index)的期刊交互引用和共被引資料為例,共計7121筆期刊,期間的交互引用資料超過1624萬筆,以這些資料計算上述的10種相關性測量方式,並且應用VxOrd進行視覺化計算。研究結果發現:在各種覆蓋率之下,交互引用的cosine(IC-Cosine)和以cosine為基礎的K50(IC-K50)兩種測量方式比其他的測量方式在預測分類時較為準確,相較於需要較多計算資源的Pearson相關係數在應用上較為可行。不論是交互引用或是共被引資料的原始次數在這個利用分類做為準確率評估標準的研究裡,都不理想。此外,交互引用的各種測量方式大多比共被引資料的測量方式更為準確。並且最為特別的是經過VxOrd的視覺化處理,各種測量方式都有比原先的測量方式得到更高的準確率。

The authors propose a new framework for assessing the performance of relatedness measures and visualization algorithms that contains four factors: accuracy, coverage, scalability, and robustness.

This method was applied to 10 measures of journal–journal relatedness to determine the best measure. The 10 relatedness measures were then used as inputs to a visualization algorithm to create an additional 10 measures of journal–journal relatedness based on the distances between pairs of journals in two-dimensional space. This second step determines robustness (i.e., which measure remains best after dimension reduction).

Results show that, for low coverage (under 50%), the Pearson correlation is the most accurate raw relatedness measure. However, the best overall measure, both at high coverage, and after dimension reduction, is the cosine index or a modified cosine index. Results also showed that the visualization algorithm increased local accuracy for most measures.

The two main groups of measures are intercitation measures, or those based on one journal citing another, and cocitation measures, which are based on the number of times two journals are listed together in a set of reference lists.

Although raw frequency has been used for both journal citation (Boyack, Wylie, & Davidson, 2002) and journal cocitation analysis studies in the past (McCain, 1991), it is rarely used today.

For intercitation studies, normalized frequencies such as the cosine, Jaccard, Dice, or Ochiai indexes (Bassecoulard & Zitt, 1999) are very simple to calculate, and give much better results than raw frequencies (Gmur, 2003).

A new type of normalized frequency, specific to journals, has been proposed recently (Pudovkin & Fuseler, 1995; Pudovkin & Garfield, 2002). This new relatedness factor (RF), an intercitation measure, is unique in that it is designed to account for varying journal sizes, thus giving a more semantic or topic-oriented relatedness than other measures.

The Pearson correlation coefficient, known as Pearson’s r, is a commonly used measure for journal intercitation (Leydesdorff, 2004a, 2004b), journal cocitation (Ding, Chowdhury, & Foo, 2000; McCain, 1992, 1998; Morris & McCain, 1998; Tsay, Xu, & Wu, 2003), document cocitation (Chen, Cribbin, Macredie, & Morar, 2002; Gmur, 2003; Small, 1999; Small, Sweeney, & Greenlee, 1985), and author cocitation studies (cf. White, 2003; White & McCain, 1998).

Lists of relatedness measurements are rarely analyzed directly, but are used as input to an algorithm that reduces the dimensionality of the data, and arranges the tokens on a 2-D plane. The distance between any two tokens on the 2-D plane is thus a secondary (or reduced) measure of relatedness.

Validation of relatedness measures has received little attention over the years. Most of these efforts have been to compare 2-D maps obtained from MDS with some sort of expert perceptions of the subject field.

Only one study has compared citation-based relatedness measures. Gmur (2003) compared six different relatedness measures based on the cocitation counts of 194 highly cited documents in the field of organization science. The measures included raw frequency, three forms of normalized frequency, Pearson’s r, and loadings from factor analysis. The bases for comparison were network-related metrics such as cluster numbers, sizes, densities, and differentiation. Results were strongly influenced by similarity type. For optimum definition of the different areas of research within a field, and their relationships, clustering based on Pearson’s r or on the combination of two types of normalized frequency worked best.

Accuracy refers to the ability of a relatedness measure to identify correctly whether tokens (e.g., journals, documents, authors, or words) are related.

Local accuracy refers to the tendency of the nearest tokens to be correctly placed or ranked. Ideally, local accuracy is measured from the perspective of each individual token. For authors, the question might be whether an author would agree with the ranking of the 10 most closely related authors. For journals, the question might be whether the closest journals were in the same discipline. For papers, the question might be whether the closest papers were on the same topic.

Global accuracy refers to the tendency for groups of tokens to be correctly placed or ranked, and requires that the tokens be clustered.

The assessment of accuracy requires some sort of independent data to use as a basis of comparison.

Coverage helps to assess the impact of thresholds on accuracy. In this analysis, thresholds are used to identify all relationships that are at or above a certain level of accuracy. Very high thresholds of relatedness will tend to identify the relationship between a few tokens, lower thresholds will include more tokens, but the level of accuracy will likely be lower.

Scalability refers to the ability of a measure (or a derived measure from a visualization program) to be applied to extremely large databases.

Robustness refers to the ability of a measure to remain accurate when subjected to visualization algorithms. Visualization algorithms reduce the dimensionality of the data, and it is reasonable to assume that the reduction in dimensionality will affect the accuracy of the measure. While the visualizations allow a user to gain insights into the underlying structure of the data, these insights should be qualified by an assessment of the concurrent loss of accuracy.

One expectation is that greater coverage will result in lower accuracy.

Another expectation is that the measures that utilize more data and more calculations will be more accurate but less scalable.

A third expectation is that accuracy will drop when a measure is subjected to dimension-reduction techniques because the underlying data is inherently multidimensional.

The last tradeoff refers to the choice of intercitation versus cocitation measures. On the one hand, intercitation-based measures should be more accurate because the data are more current (current year to past years rather than past-year pairs). On the other hand, cocitation measures can cover far more sources.

The data used to calculate relatedness measures for this study were based on intercitation and cocitation frequencies obtained from the ISI annual file for the year 2000. Science Citation Index Expanded (SCIE; Thomson ISI, 2001a) and Social Science Citation Index (SSCI; Thomson ISI, 2001b) data files were merged, resulting in 1.058 million records from 7349 separate journals. Of the 7349 journals, we limited our analysis to the 7121 journals that appeared as both citing and cited journals. There were a total of 16.24 million references between pairs of the 7121 journals.

The resulting journal–journal citation frequency matrix was extremely sparse (98.6% of the matrix has zeros). While there was a great deal more cocitation frequency information, the journal–journal cocitation frequency matrix was also sparse (93.6% of the matrix has zeros).

The 10 relatedness measures used in this study are given below, along with their equations. The six intercitation measures are raw frequency, Cosine, Jaccard, Pearson’s r, the recently introduced average relatedness factor of Pudovkin and Garfield (2002), and a new normalized frequency measure that we introduce here, K50. ... Note that the new measure, K50, is simply the cosine index minus an expected cosine value. ... The four cocitation measures are raw frequency, cosine, Pearson’s r, and the cocitation version of the K50 measure.

As mentioned above, for each of the 10 relatedness measures, a dimension reduction was done using VxOrd. The process for calculating “re-estimated measures” is as follows. First, 2-D coordinates were calculated for each of the 7121 journals using VxOrd (cf. Figure 2). Next, the distances between each pair of journals (on the 2-D plane) were calculated for the entire set and used as the re-estimated measures of relatedness.

The IC-Pearson measure is the most accurate for higher absolute levels of relatedness (up to a rank of ~85,000). As ranked relatedness increases, the curves for all but the IC-Raw measure converge. IC-Cosine, IC-K50, and IC-Jaccard measures generate nearly identical results over the entire relatedness range
up to a rank of ~125,000.

The CC-Pearson measure is the best of the four up to a rank of ~350,000, and then
drops below the CC-Cosine and CC-K50. The CC-K50 is slightly more accurate than the CC-Cosine, and the raw frequency measure, CC-Raw, gives the worst results by far.

Figure 4a shows that for the intercitation measures, the IC-Cosine and IC-K50 measures cover more journals than the other measures over the entire range of rank relatedness. The IC-Jaccard and IC-RFavg measures have the next highest coverage, followed by the IC-Pearson. The IC-Raw covers the fewest journals over most of the range.

The CC-Cosine and CC-K50 have the highest coverage, followed by the CC-Pearson. Once again, raw frequency gives the worst results.

The IC-Pearson measure is more accurate for up to a coverage of 0.58, while the IC-Cosine and IC-K50 are more accurate for coverage past 0.58. Note that, excepting the raw frequency measures, both of which do poorly, the intercitation measures are more accurate than the cocitation measures.

First, the IC-Cosine, IC-K50, and IC-Jaccard measures all have roughly comparable accuracy over the entire range of coverage. The IC-K50 measure is slightly more accurate than the others from 20–50% coverage, while the IC-Cosine is the most accurate from 50–90% coverage. The IC-Pearson measure remains below these three over the entire coverage range.

Second, the intercitation measures are more accurate than the cocitation measures in all cases.

Third, the Pearson measures are less accurate than the cosine measures for both the intercitation and cocitation data.

Also, note that the re-estimated K50 measures are essentially identical to the cosine measures for both the intercitation and cocitation data. Any differences at a particular coverage value are small enough to justify using the cosine value, which requires less calculation. It appears that, although the K50, by virtue of subtracting out the expected values, gives different individual similarity values and rankings, the aggregate effect on overall accuracy is minimal.

The most striking result comes from a comparison of the results of Figures 5 and 6, namely that the overall accuracy for all re-estimated measures is higher than for the raw measures over nearly the entire coverage range. This is an extremely counterintuitive finding, given the prevailing and common belief that information is lost when dimensionality is reduced.

Three of the intercitation measures (IC-Cosine, IC-K50, and IC-Jaccard) perform similarly, all with high-accuracy values at the both the 50% and 95% coverage levels.

All of the intercitation measures are limited to use within the citing journal set. If coverage outside the citing journal set is desired, cocitation measures can be used. Of these, the new measure introduced in this paper, CC-K50, is slightly better than the Cosine at high-coverage levels. Both the CC-Cosine and CC-K50 are clearly better than the Pearson correlation, both in terms of accuracy, and in that they do not require n(square)  calculations, and thus scale to much larger sets than the Pearson.

First, we expected the Pearson correlation to provide the best results. The reason for this expectation is that the Pearson correlation uses more information in its construction (nearly the entire intercitation or cocitation matrix) than do the other measures. Pearson correlations allow for the influence of other parties. On the other hand, the other measures only use a small amount of the data in the matrix, and tend to limit their focus to the relationship between the two journals in
question.

The second surprise was the increase in performance from the visualization software. We expected the performance to deteriorate due to the simple rule of thumb that reducing data to two dimensions requires tradeoffs that would result in lower accuracy.

The improvement in performance may be explained by the peculiarities of the VxOrd force directed algorithm. VxOrd balances attractive forces between nodes (the similarity values) with those of a repulsive grid that tries to force all nodes apart. It also cuts edges once the similarity-to-distance ratio falls below a threshold, and in most cases cuts about 50% of the original edges, thus leaving edges only where particularly strong similarities exist among a set of nodes. These dominant similarities are likely to be very accurate on the whole, and when concentrated by pruning the less accurate edges, may increase the overall accuracy of the solution.