顯示具有 clustering 標籤的文章。 顯示所有文章
顯示具有 clustering 標籤的文章。 顯示所有文章

2017年11月16日 星期四

Gowanlock, M., & Gazan, R. (2013). Assessing researcher interdisciplinarity: A case study of the University of Hawaii NASA Astrobiology Institute. Scientometrics, 94(1), 133-161.

Gowanlock, M., & Gazan, R. (2013). Assessing researcher interdisciplinarity: A case study of the University of Hawaii NASA Astrobiology Institute. Scientometrics94(1), 133-161.

本研究結合書目計量學技術與機器學習演算法評估夏威夷大學NASA天體生物學研究中心(UHNAI)的研究跨學科性(interdisciplinarity of research)。以UHNAI發表的論文資料為研究的單位,彙整論文本身的摘要和其引用的參考文獻的摘要代表該論文,UHNAI團隊共計有731篇論文,這些論文引用的相關摘要則有10216筆,另外根據其引用的期刊主題分類(Journal Subject Categories)分布製作13個合併主題分類(conflated SCs)。探討的主要問題包括:(1)評估WoK的期刊主題分類是否適合用來標註天體生物學的文獻;(2)利用合併的主題分類確認天體生物學實際與潛在的跨學科研究案例;(3)利用代表UHNAI團隊研究軌跡(research tracks)的彙整摘要確認天體生物學實際與潛在的跨學科研究案例,並確認研究人員間的潛在合作機會。

本研究採用Porter et al. (2007)依據美國國家科學院(National Academies)2005年對跨學科研究(interdisciplinary research)和多學科研究(multidisciplinary research)的定義,前者係整合來自兩個或以上學科的概念、理論、技術與(或)資料,後者則不須跨學科整合,僅需要採用其他專業知識體系的元素,導引出大於各部分總和的研究。

根據van Leeuwen (2007)的整理,研究的跨學科性可分為兩種書目計量取向,一為考慮較高層次的出版物聚合,例如國家或大學等的研究產出,通常利用現成的主題分類(SCs),將它視為是學科界限(disciplinary boundaries),進行由上而下(top-down)的分析,例如van Raan and van Leeuwen (2002) 和 Porter et al. (2007)利用SCs做為測量作者、期刊或研究領域跨越科學領域的基準;另一為單一文獻及其引用論文的由下而上(bottom-up)分析,根據作者在文件題名、摘要、關鍵詞或全文上的詞語,描述研究人員,期刊或整個領域的結構,並建議有效的未來方向。由於本研究關注未來的整合,而不是過去的產出,因此採用後者的方式。過去這方面的研究有 Kostoff et al., (2001)以引用與被引用論文的文字欄位,進行片語(phrase)的頻率和集群分析,了解研究的影響力和跨學科性;以及Rafols and Meyer (2010)結合由上而下與由下而上的方式測量學科多樣性與知識的整合。

另外,在分析跨學科邊界的合作可能性時,一般會去採訪領域專家。但這種方法受到樣本大小與主觀性等限制(Zhang et al., 2011),另外分析像是天體生物學這樣跨越多個學科的主題時,需要相當博學專家的知識。在考慮上述的限制後,本研究建議天體生物學的科際整合性分析,應由一或多個精通該主題的人員指導,但其專業知識不必要涵蓋所有的組成學科。因此,本研究採用不需要先備知識的非監督式方法來發現資料的趨勢,該方法為(Slonim et al., 2002)提出的sIB (sequential Information Bottleneck)文本集群分析。透過文本集群和分類可以描述合作與知識整合等現象,對於研究軌跡底層結構的揭示能夠提供對天體生物學研究者有用的結論。

本研究的結果顯示經由對摘要資料進行文本探勘產生的集群通常與SCs不太一致。 因此,本研究認為SCs不太適合應用於天體生物學出版物的分類,並且推測其他科際整合領域也是如此。其原因一個解釋是天體生物學研究成果會引用單學科和跨學科的出版物,可能會阻止SCs形成凝聚性的集群。此外,正如Small(2010)中所討論的,許多期刊發表了高度多樣化的內容,以期刊做為分類層級的系統並不能完全表現。10個集群是最合適的分類結果。太少的集群,無法表現來源文件的學科整合多樣性;過多的集群,則可能會太分散,從而減少發現來自不同學科和主題的共同性的機會。本研究建議,當來自不同SCs的文件聚集在一起時,這可能表明隱含的跨學科聯繫,某一個領域的知識可能對另一個領域有啟發的效用。由組成學科的研究人員評估這些共同的文件,可能可以提供一個科際整合科學發生的機制,並提供了一個潛在的跨學科合作的起點。而且利用sIB進行彙整摘要資料的文本探勘也適合用於發現合作的機會,本研究發現來自相同學術部門的作者的論文較有可能集群在一起,這也證實本研究使用的方法所產生的集群內能聚集相似的論文。因此,論文在同一集群下的作者可能可以進行生產性較高的合作。而論文分散在多個集群的作者可能表示他們有參與跨學科研究。研究也發現UHNAI的研究人員和博士後研究人員的論文大部分出現在多個集群中,是跨學科研究的主要族群。

In this study, we combine bibliometric techniques with a machine learning algorithm, the sequential Information Bottleneck, to assess the interdisciplinarity of research produced by the University of Hawaii NASA Astrobiology Institute (UHNAI).

In particular, we cluster abstract data to evaluate Thomson Reuters Web of Knowledge subject categories as descriptive labels for astrobiology documents, assess individual researcher interdisciplinarity, and determine where collaboration opportunities might occur.

Following van Leeuwen (2007), we distinguish between a top-down bibliometric approach, where large-scale trends at the highest levels of publication aggregation are considered (such as the research output of a country or university), and prefer a bottom-up approach, where we analyze individual documents and the papers they cite.

A common method used to examine the potential of collaboration across disciplinary boundaries is to interview domain experts, but this method suffers from several limitations, such as sample size and subjectivity problems (Zhang et al., 2011). Furthermore, given that the subject matter of astrobiology spans many disciplines, meaningful analysis of the responses would require the knowledge of an astrobiology polymath.

After considering these limitations, we suggest that measuring interdisciplinarity should be guided by one or more individuals versed in astrobiology, but whose expertise need not span all of its constituent disciplines. Therefore, an unsupervised approach is optimal as such methods can find trends in data without prior knowledge of its structure.

In this pilot study, we investigate the use of an unsupervised machine learning clustering technique, the sequential Information Bottleneck (sIB) (Slonim et al., 2002) to aid in measuring researcher interdisciplinarity.

Furthermore, we assess the extent to which Journal Subject Categories from the Thomson Reuters Web of Knowledge database suite are sufficient for labelling astrobiology documents.

The clustering and classification of text allow interdisciplinary analysis that 1) describes collaboration and the integration of knowledge and 2) draws conclusions that are useful to astrobiology researchers by uncovering the underlying structure of research tracks.

The multidisciplinary context given by astrobiology affords an excellent opportunity to examine the methods used to study researcher interdisciplinarity and knowledge integration.

Furthermore, we propose an iterative process to identify specific publications that bridge diverse fields, to facilitate interdisciplinary collaborations and ease the cognitive load of a single researcher who wishes to integrate knowledge from multiple disciplines.

Research that occurs at the intersection between disciplines is thought to lead to great advances in science (Porter and Rafols, 2009).

We adopt the definition suggested by Porter et al. (2007), which followed the definition given by the National Academies (2005): interdisciplinary research requires an integration of concepts, theories, techniques and/or data from two or more bodies of specialized knowledge. Multidisciplinary research may incorporate elements of other bodies of specialized knowledge, but without interdisciplinary synthesis (Wagner et al., 2011) that leads to research that is greater than the sum of its parts.

The usefulness of bibliometric indicators depends critically on the level at which we wish to understand the integrative process. For example, funding agencies may only require high-level publication co-authorship and collaboration statistics, describing the research performed by their grantees and the diversity of their home disciplines, but not addressing the essential aspect of synthesis.

Top-down approaches have been used to map scientific literature (for example, see Boyack et al. (2005)), and often represent broad areas of science with Web of Knowledge (WoK) subject categories (SCs). For example, van Raan and van Leeuwen (2002) and Porter et al. (2007) used SCs in their methodology to measure interdisciplinarity. In these studies, SCs have been employed as de facto disciplinary boundaries, and as a benchmark to measure how much a given author, journal or research area crosses scientific fields.

Unfortunately, low-level conclusions that might inform potentially productive individual collaborations cannot be made when relying on these top-down approaches, as they focus on past outputs rather than future integration.

Conversely, bottom-up bibliometric approaches incorporate the authors’ own words, in free-text fields such as: titles, abstracts, keywords1 and the full text of a document. Clustering bibliometric data at this level can describe the structure of a researcher, journal or an entire field, and suggest productive future directions.

A study by Rafols and Meyer (2010) combines bottom-up and top-down approches to measure both disciplinary diversity and knowledge integration.

While bibliometric studies tend to rely on a citation analysis, such an analysis may not be appropriate for every discipline or field. For example, a given field may tend to reference conference proceedings, websites, newspapers, or colloquia which are not as conducive to a co-citation analysis as journal articles. Due to this observation, Sugimoto (2011) suggests that studying interdisciplinarity should include publications beyond journal articles.

One of the goals of this research is to uncover the underlying structure within an astrobiology research team that undertakes interdisciplinary projects at the macro scale, but may differ in the extent of interdisciplinary work at the micro level.

To understand the research structure, we examine the abstract text of research publications and employ a method from the field of information theory, the sIB method, to cluster our high dimensional abstract data.

An advantage of using WoK for bibliometric studies is that it provides a mapping of SCs to each journal. Given the incommensurability of other bibliometric data (for example, journals do not agree upon a common set of keywords), SCs provide a way to compare publications on the journal level.

In Porter et al. (2007), the authors examine the references in sets of journal articles gathered from WoK, and relate the journals to their corresponding SCs. In this approach, a more diverse set of SCs that represent a paper derived from its references indicates a higher degree of interdisciplinarity than a set of similar SCs that represent a paper.

In particular, we combine all of the abstracts of all of the references cited by a UHNAI publication, and use these aggregated abstracts to represent each publication.

In another text mining study (Kostoff et al., 2001), employed free-text fields (such as title, keywords and abstracts) of cited/citing publications in combination with phrase frequency analysis and phrase clustering analysis to obtain a low-level understanding of research impact and interdisciplinary research.

In the following subsections, we describe our methods used to achieve the following goals:
• Examine whether WoK SCs are sufficient for labelling astrobiology documents.
• Identify actual and potential instances of interdisciplinary research in astrobiology using conflated SCs (Section 3.3).
• Identify actual and potential instances of interdisciplinary research and identify potential collaboration opportunities between researchers using aggregated abstracts to represent the research tracks of the UHNAI team (Section 3.4).

We chose this clustering method over others because it has been shown to perform better than other unsupervised clustering methods, such as k-means (Slonim et al., 2002). Furthermore, the approach should allow us to identify instances of interdisciplinary research by examining the cluster membership of our abstract data without prior knowledge of the data’s properties. It is necessary to use an unsupervised clustering method because a canonical set of astrobiology documents with which to train a clustering technique does not exist.

We modify the SCs using the following method:
• Journals with a single WoK SC that appears 10 or more times in our dataset uses the assigned WoK SC name.
• Journals with a single WoK SC that appears less than 10 times is changed to a broader WoK category (e.g. “Biochemical Research Methods” becomes “Biochemistry & Molecular Biology”).
• Journals with two or more SCs of roughly equivalent weight are assigned a new conflated SC (e.g. “Astrophysics & Geophysics”).
• Journals with two or more SCs that have a clear primary SC have “-Multidisciplinary” appended to the primary name.

The dataset has 10216 abstracts integrated over 13 conflated SCs.

We use the Synthetic Minority Over-sampling Technique (SMOTE) (Chawla et al., 2002) to produce synthetic feature vectors, where a feature vector (or feature) is a normalized numerical representation of the words that describe each abstract/instance.

We use SMOTE to create synthetic feature vectors for the minority SCs such that each SC is represented by the same number of features.

The dataset contains 731 publications by the UHNAI team.

Each publication is represented by its own abstract and the abstract of each cited publication. We aggregate all of these abstracts in a single feature vector to represent each UHNAI publication. Non-journal publications such as book chapters, conference proceedings and dissertations were included in the dataset, although they constitute a very small fraction of the total publications.

For the purposes of this paper, where our goal is to identify actual and potential instances of interdisciplinary research in astrobiology, a meaningful cluster relationship is one where papers from two or more SCs cluster together, or when researchers from different fields have the aggregated abstracts of their papers cluster together.

We begin by estimating the extent to which conflated Web of Knowledge Subject Categories accurately describe the content of astrobiology publications.

However, when abstracts are assigned one of five clusters (Figure 4-top panel), we observe that the cluster membership for most SCs is heterogeneous: there is no clear correspondence between a cluster and a single dominant SC. Even the most common SC, Astronomy & Astrophysics, is primarily distributed across the first three clusters, but is represented in all five.

Table 5 does suggest some areas in which SCs may be more appropriate document labels. For example, Oceanography appears in only one cluster, and the Multidisciplinary Sciences SC is fairly evenly distributed across four of the five. However, when increasing the number of clusters to 10, 15, and 20 (Figure 4), the heterogeneity of SCs within an individual cluster becomes even more pronounced.

One would intuitively expect more SC heterogeneity within each cluster; however, increasing the number of clusters also allows more potential of each SC to dominate a single cluster. When we increase the number of clusters to 10 (Figure 6, Table 7), we find that most of the SCs disperse into multiple clusters.

One way to interpret this result is that more clusters allow finer distinctions between content to be revealed.

At the 10 cluster level, more clusters contain single dominant SCs than at the 5 cluster level.

Figure 7 (Table 8) and Figure 8 (Table 9) present the results of clustering the abstracts into 15 and 20 clusters, respectively. We observe that many of the SCs are found distributed in multiple clusters.

Therefore, at these clustering levels, we operationalize a dominant SC within a cluster as one that either constitutes 50% or more of the abstracts alone, or one that is within 50% of the size of the most common SC5 .

By this approximation, the results at the 10 cluster level hold: as a group, the Biochemistry and Biotechnology-related SCs dominate the fewest clusters; the Astronomy, Oceanography and Physics group slightly more, and the Geochemistry and Geophysics SCs are again the most diverse, short of the Multidisciplinary Sciences SC.

Overall, at the 10 cluster level, more clusters contain single dominant SCs than at the 5, 15 or 20 cluster levels, and the usefulness of SCs as document labels reaches a relative maximum.

In some cases, the trial processes reveal some inconsistencies in the cluster membership of SCs.

Certain related SCs tend to consistently cluster together, which suggests that SCs are sufficient for characterizing astrobiology publications. However, other SCs have a limited effectiveness as document labels in this interdisciplinary domain, as some SCs did not map well to successively smaller cluster sizes.

Therefore, our results suggest that WoK SCs may not consistently reflect the diverse content of astrobiology publications.

Across all three trials at the 10 cluster level in Figure 6, a single clearly dominant SC could be identified in 27 of the 30 clusters. The Astronomy, Oceanography and Physics SCs demonstrated somewhat less monodisciplinary dominance at the 10 cluster level; all had roughly 20% of their abstracts assigned to other clusters. The Geochemistry & Geophysics and Environmental Sciences SCs demonstrated the most diversity apart from the pure Multidisciplinary Sciences SC, though somewhat surprisingly, the Geochemistry & Geophysics-Multidisciplinary SC appeared in fewer clusters than its core SC.

Analyzing the heterogeneous cluster membership of publications from diverse SCs is one way to assess interdisciplinary research possibilities, but the probabilistic nature of this method should be emphasized. A heterogeneous cluster could indicate that SCs are poor document labels, or that the clustering level should be adjusted to better match the data and metadata, or that a potential interdisciplinary relationship exists. In either case, this process could inform targeted, iterative investigation.

This result suggests that the sIB technique is able to cluster similar research on a high-level; however, utilizing more clusters should provide a lower-level view of overlap in research interests between the authors

When running the sIB technique for 10 clusters, we begin to see where researchers may find potential collaboration opportunities, and we observe which authors have specialized or broad research interests. Research can be specialized but still integrate methods, techniques and data from multiple disciplines. We believe that an author who is represented primarily in a single cluster may not be engaging frequently in interdisciplinary research, or may be focusing on narrow research problems, or using similar research methods or equipment. In Figure 10, we see that the two astrochemists (Bennett and Kaiser) are entirely represented by cluster 8, consistent with the results presented in Figure 9. We know that their research is heavily influenced by their experimental apparati, thus suggesting that the experimental methods and apparati significantly affect the description of a research track. Interestingly, Sch¨orghofer’s research is on various planetary bodies such as Mars and the Moon, which is also true of Taylor. Therefore, clustering the text of the aggregated abstracts sufficiently illuminates similarities in research tracks across disciplinary boundaries, in this case, between astronomy and geology.

In Figure 11, we observe that Huss, Jewitt, Krot and Meech’s research is found in many clusters. This signifies that their research is likely to be very interdisciplinary. With regards to those authors represented by a few clusters, we cannot conclude that their research is absolutely mono-disciplinary, as it may be very specialized, or utilize the same methods or apparati. However, we believe that those UHNAI authors with publications in multiple clusters are more likely to be engaged in interdisciplinary research. In Figure 12, we observe that of the senior (non-postdoctoral fellows) astronomers (Reipurth, Meech, Jewitt, Haghighipour, Owen, Sch¨orghofer) half (Meech, Jewitt, and Owen) are fairly diverse in their research interests and the other half (Reipurth, Haghighipour, Sch¨orghofer) are engaged in specialized or mono-disciplinary research.

These results suggest that the sIB method, in combination with aggregated abstracts, can illuminate areas of implicit commonality where the research areas of scientists from diverse disciplines overlap. Furthermore, while clusters do not inherently relate any information about a researcher’s discipline, it is clear that researchers from the same department often cluster together. Therefore, we expect that performing a similar analysis on the entire NASA Astrobiology Institute will show where collaborations between researchers can occur, and can assist NASA with outlining research priorities. These results can serve as the framework for a geospatial visualization of common yet unconnected research tracks and potential collaborators, similar to the “hot regions” described by Bornmann and Waltman (2011).

2016年7月10日 星期日

Janssens, F., Zhang, L., De Moor, B., & Glänzel, W. (2009). Hybrid clustering for validation and improvement of subject-classification schemes. Information Processing & Management, 45(6), 683-702.

Janssens, F., Zhang, L., De Moor, B., & Glänzel, W. (2009). Hybrid clustering for validation and improvement of subject-classification schemes. Information Processing & Management45(6), 683-702.



科學認知映射(cognitive mapping of science)能將科學的結構(the structure of science)加以視覺化,起初應用於資訊服務(information services)、後來發現也可將其應用在科學政策(science policy)與研究評鑑(research evaluation),現在則有越來越多將其應用在發現新興與正在融合中的領域以及主題劃分(subject delineation)的改善。

科學認知映射主要可分為依據引用資訊、依據文本以及混合上述兩種資訊等三種方法。本研究利用文本與引用資訊混合的方法,將2002-2006年Web of Science資料庫內的期刊進行集群,以集群結果產生的認知映射,檢驗目前的期刊主題分類架構,如果可行的話,也將提出改善方式。

本研究首先評估ESI (Essential Science Indicators) 的22個領域主題分類架構,並且將其視覺化。圖1的左右分別是以交互引用與文本方式測量22個ESI領域的Silhouette值,從圖上發現生物學及生物化學(#2)、臨床醫學(#4)、工程學(#7)、植物及動物科學(#19)以及社會科學(#21)等領域上的期刊並沒有足夠好的一致性(coherent)。



從詞語的TF-IDF可以找出每個領域的描述詞語,而這裡也可發現不少領域的描述語互有重疊,例如工程學(#7)與電腦科學(#5)、化學(#3)與材料科學(#11)、植物及動物科學(#19)與環境/生態學(#8),以及生物學及生物化學(#2)、分子生物學及遺傳學(#14)與臨床醫學(#4)等等,此外,從描述社會科學(#21)的詞語也可以了解這個領域高度的異質性(heterogeneity)。


圖2則是以Pajek畫出22個ESI領域的結構圖,圖上也可發現生物學及生物化學(#2)與臨床醫學(#4)、化學(#3)與材料科學(#11)、電腦科學(#5)與工程學(#7)、環境/生態學(#8)與植物及動物科學(#19)等領域之間有很強的關連。

然後將約8300種期刊利用餘弦相似法及Wade的凝聚式階層集群演算法(Wade's  agglomerative hierarchical cluster algorithm)進行集群,再比較集群結果與分類架構。值得說明的是文字部分的資訊可以提供集群結果的標示(labelling),而引用部分則可產生交互引用圖(cross-citation graph)提供視覺化,並且輸入PageRank演算法以決定代表性期刊。

決定集群的數目可以根據集群結果的品質,而集群品質有內在或外在驗證測量(internal or external validation measures)等兩種評估方式。內在驗證只考慮資料與集群的統計特性,例如dendrogram、Silhouette值與模組性(modularity)等;外在驗證需要將集群結果與一個已知的劃分標準進行比較,例如計算兩者間的Jaccard相似性 (Jaccard similarity)。本研究以dendrogram的視覺化方式將期刊首先分為三大群,再分為七群,最後分為22群,三大群約等於自然與應用科學(生物學、農學及環境科學;物理、化學及工程學;數學及電腦科學)、醫學與社會科學以及人文學。以TF-IDF描述語來看,七個群組中有三個屬於自然與應用科學,兩個是生命科學(生物科學及生物醫學與臨床、實驗醫學及神經科學)與兩個是社會科學以及人文學(經濟學、商學及政治學與心理學、社會學及教育學)

圖6是22個群組的結構圖,圖上可看到屬於社會科學以及人文學的群組(#1、#6、#14及#22與#9、#11及#21)、地理學、環境科學、生物學及農學(#2、#15及#19)、物理、化學及工程學(#4、#20及 #5)、數學及電腦科學(#8及#18)、生物科學及生物醫學(#3、#13及#16)與臨床、實驗醫學及神經科學(#7、#10、#12及#17)。



表3比較22個ESI領域與本研究利用引用、文本以及混合等方法產生的22個群組的集群品質,可以發現混合引用及文本資料在各種指標上幾乎都有最好的表現。


圖8是利用Jaccard指標比較集群結果與ESI架構的一致性(concordance)。



最後,本研究並分析期刊轉移(migration)的情形,也就是期刊不屬於原本ESI架構的領域所對應的集群,而被分配到另一個不同集群的現象,好的轉移(Good migration)能使分類的一致性增加,也就是Silhouette值或是模組性增加。本研究希望利用這個現象,從集群與領域的一致性的基礎上提出改善目前期刊主題分類架構的方法。

The main bibliometric techniques are characterised by three major approaches, particularly the analysis of citation links (cross-citations, bibliographic coupling, co-citations), the lexical approach (text mining), and their combination.

A hybrid text/citation-based method is used to cluster journals covered by the Web of Science database in the period 2002–2006. The objective is to use this clustering to validate and, if possible, to improve existing journal-based subject-classification schemes.

In a first step, the 22-field subject-classification scheme of the Essential Science Indicators (ESI) is evaluated and visualised. In a second step, the hybrid clustering method is applied to classify the about 8300 journals meeting the selection criteria concerning continuity, size and impact.

The hybrid method proves superior to its two components when applied separately. The choice of 22 clusters also allows a direct field-to-cluster comparison, and we substantiate that the science areas resulting from cluster analysis form a more coherent structure than the ‘‘intellectual” reference scheme, the ESI subject scheme.

Moreover, the textual component of the hybrid method allows labelling the clusters using cognitive characteristics, while the citation component allows visualising the cross-citation graph and determining representative journals suggested by the PageRank algorithm.

Finally, the analysis of journal ‘migration’ allows the improvement of existing classification schemes on the basis of the concordance between fields and clusters.

The history of cognitive mapping of science is as long as the history of computerised scientometrics itself. While the first visualisations of the structure of science were considered part of information services, i.e., an extension of scientific review literature (Garfield, 1975, 1988), bibliometricians soon recognised the potential value of structural science studies for science policy and research evaluation as well. At present, the identification of emerging and converging fields and the improvement of subject delineation are in the foreground.

The main bibliometric techniques are characterised by three major approaches, particularly the analysis of citation links (cross-citations, bibliographic coupling, co-citations), the lexical approach (text mining), and their combination.

For instance, clustering based on co-citation and bibliographic coupling has to cope with several severe methodological problems. This has been reported, among others by Hicks (1987) in the context of cocitation analysis and by Janssens, Glänzel, and De Moor (2008) with regard to bibliographic coupling. One promising solution is to combine these techniques with other methods such as text mining (e.g., combined co-citation and word analysis: Braam, Moed, & Van Raan, 1991a; combination of coupling and co-word analysis: Small (1998); hybrid coupling-lexical approach: Janssens, Glänzel, & De Moor, 2007; Janssens et al., 2008).

Jarneving (2005) proposed a combination of bibliometric structure–analytical techniques with statistical methods to generate and visualise subject coherent and meaningful clusters. His conclusions drawn from the comparison with ‘intellectual’ classification were rather sceptical.

Despite several limitations, which will be discussed further in the course of the present study, cognitive maps proved useful tools in visualising the structure of science and can be used to adjust existing subject-classification schemes even on the large scale as we will demonstrate in the following.

The main objective of this study is to compare (hybrid) cluster techniques for cognitive mapping with traditional ‘intellectual’ subject-classifications schemes.

In a first study by authors related to the current work, the pilot study of Glenisson, Glänzel, & Persson (2005), further extended and confirmed by Glenisson, Glänzel, Janssens et al. (2005), full-text analysis and traditional bibliometric methods were serially combined to improve the efficiency of the individual methods. It was clear that clusters found through application of text mining provided additional information that could be used to extend and explain structures found by bibliometric methods, and vice versa. However, the integration was still limited to serial combination.

All textual content was indexed with the Jakarta Lucene platform (Hatcher & Gospodnetic, 2004) and encoded in the Vector Space Model using the TF-IDF weighting scheme reviewed by Baeza-Yates & Ribeiro-Neto (1999). Stop words were neglected during indexing and the Porter stemmer was applied to all remaining terms from titles, abstracts, and keyword fields. The resulting term-by-document matrix contained nine and a half million term dimensions (9,473,061), but by ignoring all tokens that occurred in one sole document, only 669,860 term dimensions were retained. Those ignored terms with document frequency equal to one are useless for clustering purposes.

The dimensionality was further reduced from 669,860 term dimensions to 200 factors by Latent Semantic Indexing (LSI) (Berry, Dumais, & O’brien, 1995; Deerwester, Dumais, Furnas, Landauer, & Harshman, 1990), which is based on the Singular Value Decomposition (SVD).

Text-based similarities were calculated as the cosine of the angle between the vector representations of two papers (Salton & Mcgill, 1986).

For simplicity and efficiency, the method used to summarise the subject of a field or cluster is based on selecting the terms with the highest mean TF-IDF weights over all journal papers in the field or cluster, where the IDF factor is calculated on the complete term-by-paper matrix (more than six million papers).

For example, Treeratpituk and Callan (2006) automatically select and assign a few concise labels to hierarchical clusters by combining statistical features from the cluster, parent cluster, and a corpus of general English into a descriptive score.

Geraci, Maggini, Pellegrini, and Sebastiani (2008) label clusters by combining intra-cluster and inter-cluster term extraction based on a variant of the information gain measure, and by looking within the titles of Web pages for the substring that best matches the selected top-scoring words.

The similarities Sij used for clustering were found by calculating the cosine of the angle between the pair of vectors containing all symmetric journal cross-citation values between the two respective journals (i and j) and all other journals (i.e., row or column of the matrix C):

The journal cross-citation graph is also analysed to identify important high-impact journals. We use the PageRank algorithm (Brin & Page, 1998) to determine representative journals in each cluster. Besides, the graph can also be used to evaluate the quality of a clustering outcome.

In order to subdivide the journal set into clusters we used the agglomerative hierarchical cluster algorithm with Ward’s method (Jain & Dubes, 1988).

In general, the number of clusters is determined by comparing the quality of different clustering solutions based on various numbers of clusters. Cluster quality can be assessed by internal or external validation measures. Internal validation solely considers the statistical properties of the data and clusters, whereas external validation compares the clustering result to a known gold standard partition

This compound strategy encompasses observation of a dendrogram, text- and citation-based mean Silhouette curves, and modularity curves. Besides, the Jaccard similarity coefficient is used to compare the obtained results with an intellectual classification scheme.

Up to a multiplicative constant, modularity measures the number of intra-cluster citations minus the expected number in an equivalent network with the same clusters but with citations given at random. Intuitively, in a good clustering there are more citations within (and fewer citations between) clusters than could be expected from random citing.

In Fig. 8, we use the Jaccard index to compare each cluster with every field from the intellectual ESI classification, in order to detect the best-matching fields for each cluster.

Nowadays two ISI systems are widely used, in particular, the ISI Subject Categories, which are available in the JCR and through journal assignment in the Web of Science as well, and the Essential Science Indicators (ESI).

While the first system assigns multiple categories to each journal and is too fine grained (254 categories) for comparison with cluster analysis, the ESI scheme is forming a partition (with practically unique journal assignment) and the 22 fields are large enough. ... This subject-classification scheme is in principle based on unique assignment; only about 0.6% of all journals were assigned to more than one field over a 5-year period.

Fig. 1 presents the evaluation of the 22 ESI fields based on the cross-citation- (left) and text-based (right) Silhouette values (see Section 3.3.3). Since the ESI fields form a partition, this approach allows to evaluate their consistency as if the fields were results of a clustering procedure. Multi-, interand cross-disciplinarity of journals can certainly affect the results.


Several fields seem not to be coherent enough from both perspectives (i.e., the cross-citation and textual approach). Above all, the Silhouette values of field #2 (Biology and Biochemistry), #4 (Clinical Medicine), #7 (Engineering), #19 (Plant and Animal Science) and #21 (Social Sciences) substantiate that at least five of the 22 fields are not sufficiently coherent.

Simultaneously to the above validation, the textual approach also provides the best TF-IDF terms – out of a vocabulary of 669,860 terms – describing the individual fields. These terms are presented in Table 2. Although these terms already provide an acceptable characterisation of the topics covered by the 22 fields, considerable overlaps are apparent between pairs of fields, respectively: Engineering (#7) and Computer Science (#5), Chemistry (#3) and Materials Science (#11), Plant and Animal Science (#19) and Environment/Ecology (#8), as well as Biology and Biochemistry (#2), Molecular Biology and Genetics (#14) and Clinical Medicine (#4). In addition, the terms characterising the social sciences (#21) reflect a pronounced heterogeneity of the field.



The structural map of the 22 ESI fields based on cross-citation links is presented in Fig. 2. For the visualisation we used Pajek (Batagelj & Mrvar, 2003). The network map confirms the strong links we have found based on the best terms between fields #2 and #14, #3 and #11, #5 and #7, and #8 and #19, respectively.

In Table 3 we compare the quality of the partition of 22 ESI fields with the quality of the 22 clusters resulting from citation-based, text-based and hybrid clustering.


The cluster dendrogram shows the structure in a hierarchical order (see Fig. 4). We visually find a first clear cut-off point at three clusters, a second one around seven, and 22 clusters also seemed to be an acceptable/appropriate number.

The number of three clusters results in an almost trivial classification. Intuitively, these three high-level clusters should comprise natural and applied sciences, medical sciences, and social sciences and humanities.

The solution comprising of seven clusters results in a non-trivial classification. The best TF-IDF terms (see Table 5) show that three of these clusters represent the natural/applied sciences, whereas two classes each stand for the life sciences and the social sciences and humanities. This situation is also reflected by the cluster dendrogram in Fig. 4. A closer look at the best TF-IDF terms reveals that the social-sciences cluster (#1 of the 3-cluster solution) is split into the cluster #1 (economics, business and political science) and #6 (psychology, sociology, education), the life-science cluster (#3 in the 3-cluster scheme) is split into clusters #3 (biosciences and biomedical research) and #7 (clinical, experimental medicine and neurosciences) and, finally, the sciences cluster #2 of the 3-cluster scheme is distributed over three clusters in the 7-cluster solution, particularly, the cluster comprising biology, agriculture and environmental sciences (#2), physics, chemistry and engineering (#4) as well as mathematics and computer science (#5).

The social-sciences and humanities clusters form two groups that are each strongly interlinked; one consists of clusters #1, #6, #14 and #22 with focus on humanities, economics, business, political and library science, the other one comprises #9, #11 and #21 with sociology, education and psychology. This is in line with the hierarchical structure shown in Fig. 4. These two groups correspond to the two social-sciences clusters in the 7-cluster solution (cf. Section 4.4).

On the basis of the most important TF-IDF terms (see Table 6) we can assign clusters #2, #15 and #19 to geosciences, environmental science, biology and agriculture, which, in turn, form a larger group corresponding to the first of the three ‘‘megaclusters” in the 7-cluster solution.

These science clusters form two groups, #4, #20 and #5 form one group of chemistry, physics and engineering, while #8 and #18 form the third group comprising mathematics and computer science.

Here we have a biomedical and a clinical group. These two groups are in line with the hierarchical structure of the dendrogram in Fig. 4 but less clearly distinguished in the graphical network presentation (Fig. 6). Nonetheless, the terms provide an excellent description for at least some of the medical clusters: cluster #7 stands for the neuro- and behavioral sciences, #3 for bioscience, #10 for the clinical and social medicine, #13 microbiology and veterinary science, #12 non-internal medicine, #16 hematology and oncology and #17 cardiovascular and respiratory medicine. According to the dendrogram clusters 3, 13, 16 and clusters 7, 10, 12, 17 form one larger cluster each. On the basis of the best terms, we can characterise these groups as the bioscience–biomedical and the clinical and neuroscience group, respectively.

In this subsection we compare the structure resulting from the hybrid clustering with the ESI subject classification. This comparison is based on the centroids of the clusters and fields. The centroid of a cluster or field is defined as the linear combination of all documents in it and is thus a vector in the same vector space. For each cluster and for each field, the centroid was calculated and the MDS of pairwise distances between all centroids is shown in Fig. 7.

In Fig. 8, we use the Jaccard index to determine the concordance between our clustering solution and the ESI Scheme by comparing each cluster with every field, in order to detect the best-matching fields for each cluster. The darker a cell in the matrix, the higher the Jaccard index, and hence the more pronounced the overlap between the corresponding cluster and ESI field.

If clustering algorithms are adjusted or changed, one can observe the following phenomenon. Some units of analysis are leaving clusters they formerly belonged to and end up in different clusters. This phenomenon is called ‘migration’. We can distinguish between ‘good migration’ and ‘bad migration’.

‘Good migration’ is observed if the goodness of the unit’s classi- fication improves, otherwise we speak about ‘bad migration’. We can also apply this notion of migration to the comparison of clustering results with any reference classification. In the following we will use the ESI scheme as reference classification.

Out of 8305 journals under study, there were more than one third, namely, 3204 journals that were not assigned to the cluster which best matches their ESI field. As already mentioned above, we call these journals ‘migrated journals’.

‘Good migrations’ are observed if journals improved their Silhouette values after migration. Based on their titles and scopes (not shown), apparently they should indeed be assigned to the cluster to which they have moved.

Although the Silhouette and modularity values substantiate a more coherent structure of the hybrid clustering as compared with the ESI subject scheme, not all clusters are of high quality. Problems have been found, for instance, in clusters #1 and #12 where interdisciplinarity and strong links with other clusters distort the intra-cluster coherence.


2016年7月7日 星期四

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

本研究利用引用文獻出現文句內容的相似性來測量作者的主題相關性(topical relatedness)。傳統的作者共被引分析(Author co-citation Analysis, ACA)做法是利用參考文獻裡被引用作者的共被引頻率(White and Griffith, 1981),然後利用Pearson相關係數 (Pearson correlation coefficient)或是 Salton提出的餘弦相似性測量作者的相似性,在書目計量學研究裡已經廣泛運用於確認與追蹤學科的知識結構(the intellectual structure of an academic discipline) (He & Hui, 2002)。然而這種做法並未考慮引用的內容,Jeong, Song, & Ding, (2014)與 Zhao & Strotmann (2014)則利用全文裡提到的作者並將有關的內容加入ACA的計算。

本研究認為累積被引作者出現的文句能夠代表作者的研究領域,因此利用JASIST的全文資料,剖析HTML,取出論文的後設資料(題名、作者姓名、出版年、DOI與摘要)、引用資訊(引用文句與參考文獻索引)以及參考資訊(作者姓名、出版年、題名與期刊)。在這個研究裡,共使用2003年1月到2015年6月的1910篇論文,合計77,408筆參考文獻。將引用文句與一般文句分開,連結文句內的參考文獻索引與參考資訊,選取100位最多被引用的作者,進行傳統的ACA以及本研究提出的新方法。本研究的新方法利用Mikolov et al., (2013)提出的Word2Vec 模型 (Word2Vec models),根據參考文獻出現的引用文句,找出作者間的相似性。Word2Vec 模型以大量的文本為基礎,利用類神經網路方法( neural network approaches),找出詞語之間的語意關係,將每一個出現於文句的詞語轉換成向量,使得這些向量之間的相似性能夠保持詞語在語意上的關係。本研究將被引用的作者姓名視為是引用文句中出現的詞語,測量作者間在研究主題的相似性與合作關係。

表2是傳統的ACA方法與本研究的方法分別找出的10組最相似的作者,本研究的方法找出10組最相似的作者中有一半是具有合作關係的作者。

另外,將兩種方法產生的作者關係分別繪製成網路圖,節點代表作者,利用PageRank決定的節點大小,節點的遠近由作者間的相似性決定,並且以Blondel, Guillaume, Lambiotte, & Lefebvre (2008)提出的社群偵測(community detection)方法進行分群。圖三與圖四分別是傳統ACA與本研究提出方法的結果。

圖三上可明顯地看到所有的作者分為兩群,依據社群偵測的分群結果,左邊的作者可再分為兩群:最左邊紅色的一群為研究資訊尋求行為(information seeking behavior)的作者,紫色的一群則與資訊檢索(information retrieval)研究有關,右邊綠色的一群則是研究書目計量學(bibliometrics)的作者。介於左右兩大群體的作者分別有兩位:Borgman和Salton。這兩位都是資訊科學領域傳統上會經常引用的作者。



在以Word2Vec方法產生的作者網路上,與資訊檢索有關的作者群組位於左方,包括上方的資訊尋求行為以及下方的文件檢索(document retrieval)兩個群組,書目計量學在圖四上則分為兩個有關的群組,一個主要包含作者分析(author analysis),另一則是期刊引用分析(journal citation analysis)與評鑑指標(evaluation indicator)。與圖三不同的是,圖四上的群組彼此間都有連結,並且圖形上更具體地呈現次學科(sub-disciplines)以及重要的作者。

Unlike other ACA studies, we used citing sentences to reflect topical relatedness of authors.

In  our  research,  we extended  traditional  approaches by  adopting Word2Vec, one  of  deep learning methods, to measure author similarity.

We also conducted in-depth network analysis of author maps.

The results of Word2Vec-based author map revealed more specific sub-disciplines and the important authors in perspective of topical influence than traditional approach does.

Author co-citation Analysis (ACA), which was introduced by White and Griffith (1981), has been widely used in bibliometrics researches to identify and trace the intellectual structure of an academic discipline (He & Hui, 2002). In ACA, traditional approaches relied on the co-citation frequency of cited authors in the reference section.

Thus, one of the main topics in ACA was methodological discussion of what kind of measure is appropriate and relevant for calculation of author similarities (Leydesdorff, 2005; van Eck & Waltman, 2007). Existing approaches based on co-citation frequencies such as Pearson correlation coefficient and Salton’s cosine similarity, however, do not capture the citation content.

Thus, some recent researches used the full-text to obtain the topical relatedness between the cited authors (Jeong, Song, & Ding, 2014; Zhao & Strotmann, 2014). They analyzed the authors mentioned in the full-text and incorporated contents related with cited authors into ACA.

In that sense, cumulated citing sentences of cited authors are able to well represent the cited researches and cited authors’ research areas. In addition, these citing sentences are particularly useful for summarization of a research document.

Figure 1 shows the overall system flow of our approach.


For content analysis, however, we collected full-text research articles of JASIST in HTML format. Through the HTML parsing process, we extracted the metadata (title, author name, year, DOI and abstract), citation information (citing sentence, and reference id), and reference information (author name, year, title, and journal).

To compare our method to traditional ACA, we computed author-pairs in both approaches. In Word2Vec-based method, the full-text data, first, are splitting into sentences. In second step, matching the citing sentences with reference id in reference section, we separated the citing sentences and other general sentences. Then, citing sentences are preprocessed in the following steps: tokenization, POS tagging, lemmatization of the tokenized sentence, and stop word removal.

From these data, we trained Word2Vec model for calculating author similarity and generated author-author similarity matrix. To compare the previous research, traditional author counting approach, we also construct co-citation matrix based on citation counts. Since we preprocessed full-text including all reference information, these matrices considered all cited authors.

To evaluation, we selected top 100 authors which are highly cited in both methodology, and conduct network analysis through visualizing author maps.

The data was gathered from 1,910 full-text articles in the JASIST digital library over 12 years (from January 2003 to June 2015). The 1,910 collected documents have 77,408 references. We extracted elements from the full-text article: 1) citing sentences from the body of the article, 2) the references information, and 3) all cited authors. Table 1 shows the basic statistics of collected data.

Word2Vec models, one of the neural network approaches, are able to carry semantic meanings and turns text into a numerical form that deep-learning nets can understand (Mikolov et al., 2013). Based on a large amount of plain text, Word2Vec trains relationships between words automatically.

Word2Vec spatially encoded a word meaning and the relationship between words, which was originally applied to word clustering or synonym detection (Wolf et al., 2014). We applied Word2Vec into author similarity measure regarding cited author names as a word in plain text.

Since authors’ oeuvre was represented as the citing sentences in research articles, the Word2Vec-based method could consider various topics of the author.

In the proposed approach, however, the author names are also trained as words in a same citing sentence. Therefore, the similarity between two authors in the Word2Vec-based method reflects both topical relatedness and collaborations.

Table 2 shows top 10 pairs by the traditional ACA method (Pearson correlation based similarity) and the Word2Vec based approach respectively. About the half of pairs resulted from the Word2Vec approach are the co-author relationship.

This results imply that the proposed approach enables to detect wider range of author pairs in perspective of topical relatedness and grasp more diverse research fields of information science.

To examine whether there are structural differences in two measures of author similarity, we constructed two author networks with top 100 authors. For network visualization, we used PageRank (Brin & Page, 1998) to determine the node size and also adopted the modularity algorithm (Blondel, Guillaume, Lambiotte, & Lefebvre, 2008) for the community detection.

Figure 3 illustrates roughly two parts that consist of information retrieval and bibliometrics, two major research areas in JASIST. The author group of information retrieval (purple) along with information seeking behavior (red) is located at the left side, and the author group related with bibliometrics is located at the right side.

There are only two authors located between two groups (Borgman and Salton), who are traditionally cited authors in the information science field. Borgman studied various topics including information retrieval and scholarly communication and wrote the important books that had won the best information science book from ASIST. Salton’s works also received a lot of citations for a long time in the field of information science.


The author group related with information retrieval in the left side of the network is split into information seeking behavior (blue) located in the upper side of the network and document retrieval (yellow) located at the bottom side of the network. The group related to bibliometrics is also separated into two parts: (1) a cluster (green) including author analysis and (2) journal citation analysis and evaluation indicator (red).

Unlike Figure 3, the communities in the network are connected to each other. Brin is connected with both document retrieval and citation analysis communities. This may be attributed to the fact that the PageRank, developed by Brin and Page (1998), is used in information retrieval and also studied in network analysis to compute node centrality.

In bibliometrics, PageRank is adopted as one of the centralities in citation networks (Ding, Yan, Frazho, & Caverlee, 2009). Ingwesen, who is located between information retrieval and bibliometrics, studied information retrieval in earlier works, he extended the research area to network analysis such as webometrics.

It implies that the authors linked by citation are topically grouped in the Word2Vec-based author network.

2015年4月9日 星期四

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

本研究比較兩種以內容為基礎的期刊分類以及兩種以演算法為基礎的期刊分類。兩種以內容為基礎的期刊分類分別是ISI的主題分類(Subject Categories)以及Glänzel and Schubert (2003)的領域/次領域分類(field/subfield classification)SOOI,兩種以演算法為基礎的期刊分類則分別是Blondel et al. (2008)提出的展開式(unfolding)社群偵測(community detection)法以及Rosvall, and Bergstrom (2008)的隨機漫步(random walk)矩陣分解(matrix decomposition)法。若是利用以內容為基礎的分類,期刊可以同時指定多個類別;以演算法為基礎的期刊分類則可以使類別內的引用(within-category citation)對類別內的引用(between-category citation)的比率最大化,也就是將期刊彼此之間的引用資料排列成矩陣,經過適當的行列排列後,使得主要對角線(principal diagonal)附近的數值較大,而其他地方則接近0。

各種分類的相關統計資料如表1所示:


由於以內容為基礎的分類方法具有多重分類特性以及以演算法為基礎的分類方法以矩陣分解為目的,從表1上可以觀察到兩種現象:1) 在類別內期刊數的中位數方面,可以看到以內容為基礎的兩種期刊分類方法較以演算法為基礎的期刊分類方法來得多,可配合圖1每個類別期刊數的分佈在0.50上所呈現的情形。另外,圖1也可發現四種分類方法都是對數常態分布(log normal distribution),也就是在這四種分類方法中,相對少數的類別擁有大量的期刊,然而許多類別卻只有少量期刊。並且以演算法為基礎的分類方法比以內容為基礎的分類方法更偏斜(more skewed),也較是上述的情況更嚴重。隨機漫步方法的前十個類別共有57%種期刊,展開方法則有50%,但ISI和SOOI則分別只有15%和31%。


2) 從引用的分布情形來看,兩種以內容為基礎的分類方法的引用次數總計比以演算法為基礎的分類方法多,但隨機漫步方法和展開方法有較多比率分布在類別內,但ISI和SOOI則是主要分布在類別之間。

接下來,以引用式樣(citation patterns)的餘弦相似性(cosine similarity),比較各種分類方法的類別彼此間的相似性。結果ISI和SOOI的中位數分別是0.020和0.066,比隨機漫步方法和展開方法的0.009和0.007高許多,其原因同樣是因為內容為基礎的方法有多重分類的特性,因此類別間的邊緣較模糊,而演算法為基礎的方法在類別間切割得較清楚。然後將各種分類方法的類別依照它們的相似性繪製成網路圖。四種方法繪製的網路圖大致上都可以看出包含兩大群,一個是生物醫學,另一個則是物理學與工程學,兩個大群體透過三個群體相連,包括化學、地理學-環境科學-生態學群體、以及電腦科學,社會科學群體在網路圖上有些分離,透過行為科學/神經科學和生物醫學相連,並且也透過電腦科學與數學和物理學/工程學相連。綜上所述,不同的科學地圖是相似的,但它們在群體內部類別的密度不同。

In this study, we test the results of two recently available algorithms for the decomposition of large matrices against two content-based classifications of journals: the ISI Subject Categories and the field/subfield classification of Glänzel and Schubert (2003).

The content-based schemes allow for the attribution of more than a single category to a journal, whereas the algorithms maximize the ratio of within-category citations over between-category citations in the aggregated category-category citation matrix.

At that time, Leydesdorff & Rafols (2009) were deeply involved in testing the ISI Subject Categories of these same journals in terms of their disciplinary organization. Using the JCR of the Science Citation Index (SCI), we found 14 major components using 172 subject categories, and 6,164 journals in 2006. Given our analytical objectives and the well-known differences in citation behaviour within the social sciences (Bensman,2008), we decided to set aside the study of the (220 − 175 = ) 45 subject categories in the social sciences for a future study.

Our findings using the SCI indicated that the ISI Subject Categories can be used for statistical mapping purposes at the global level despite being imprecise in terms of the detailed attribution of journals to the categories.

In this study, we compare the results of these two algorithms with the full set of 220 Subject Categories of the ISI. In addition to these three decompositions, a fourth classification system of journals was proposed by Glänzel and Schubert (2003) and increasingly used for evaluation purposes by the Steungroep Onderwijs and Onderzoek Indicatoren (SOOI) in Leuven, Belgium. These authors originally proposed 12 fields and 60 subfields for the SCI, and three fields and seven subfields for the Social Science Citation Index and the Arts and Humanities Citation Index. Later, one more subfield entitled “multidisciplinary sciences” was added.

Thus, because research topics are, on the one hand, thinly spread outside the core group and, on the other hand, the core groups are interwoven, one cannot expect that the aggregated journal-journal citation matrix matches one-to-one with substantive definitions of categories or that it can be decomposed in a single and unique way in relation to scientific specialties. The choice of an appropriate journal set can be considered as a local optimization problem (Leydesdorff, 2006).

Citation relations among journals are dense in discipline-specific clusters and are otherwise very sparse, to the extent of being virtually non-existent (Leydesdorff & Cozzens, 2003).

The grand matrix of aggregated journal-journal citations is so heavily structured that the mappings and analyses in terms of citation distributions have been amazingly robust despite differences in methodologies (e.g., Leydesdorff, 1987 and 2007; Tijssen, de Leeuw, & van Raan, 1987; Boyack, Klavans, & Börner, 2005; Moya-Anegón et al., 2007; Klavans & Boyack, 2009).

A decomposable matrix is a square matrix such that a rearrangement of rows and columns leaves a set of square sub-matrices on the principal diagonal and zeros everywhere else.

In the case of a nearly decomposable matrix, some zeros are replaced by relatively small nonzero numbers (Simon & Ando, 1961; Ando & Fisher, 1963). Near-decomposability is a general property of complex and evolving systems (Simon, 1973 and 2002).

The decomposition into nearly decomposable matrices has no analytical solution. However, algorithms can provide heuristic decompositions when there is no single unique correct answer.

Newman (2006a, 2006b) proposed using modularity for the decomposition of nearly decomposable matrices since modularity can be maximized as an objective function.

Blondel et al. (2008) used this function for relocating units iteratively in neighbouring clusters. Each decomposition can then be considered in terms of whether it increases the modularity.

Analogously, Rosvall, and Bergstrom (2008) maximized the probabilistic entropy between clusters by estimating the fraction of time during which every node is visited in a random walk (cf. Theil, 1972; Leydesdorff, 1991).

The data were harvested from the CD-Rom version of the JCR of the SCI and Social Science Citation Index 2006, and then combined. ... The resulting set of 7,611 journals and their citation relations otherwise precisely corresponds to the online version of the JCRs. This large data matrix of 7,611 times 7,611 citing and cited journals was stored conveniently as a Pajek (.net) file and used for further processing.

The 7,611 journals are attributed by the ISI with 11,856 subject classifiers. This is 1.56 (±0.76) classifiers per journal. The ISI staff assign the 220 ISI Subject Categories on the basis of a number of criteria including the journal's title and its citation patterns (McVeigh, personal communication, March 9, 2006; Bensman & Leydesdorff, 2009).

According to the evaluation of Pudovkin and Garfield (2002), in many fields these categories are sufficient, but the authors added that “in many areas of research these ‘classifications’ are crude and do not permit the user to quickly learn which journals are most closely related” (p. 1113).

Leydesdorff and Rafols (2009) found that the ISI Subject Categories can be used for statistical purposes—the factor analysis for example can remove the noise—but not for the detailed evaluation. In the case of interdisciplinary fields, problems of imprecise or potentially erroneous classifications can be expected.

For the purpose of developing a new classification scheme of scientific journals contained in the SCIs, Glänzel and Schubert (2003) used three successive steps for their attribution. The authors iteratively distinguished sets cognitively on the basis of expert judgements, pragmatically to retain multiple assignments within reasonable limits, and scientometrically using unambiguous core journals for the classification. The scheme of 15 fields and 68 subfields is used extensively for research evaluations by the Steunpunt Onderwijs and Onderzoek Indicatoren (SOOI), a research unit at the Catholic University in Leuven, Belgium, headed by Glänzel.

The SOOI categories cover 8,985 journals. Using the full titles of the journals, 7,485 could be matched with the 7,611 journals under study in the JCR data for 2006 (which is 98.3%). These journals are attributed 10,840 classifiers at the subfield level. This is 1.45 (±0.66) categories per journal. One category (“Philosophy and Religion”) is missing because the Arts & Humanities Citation Index is not included in our data. Thus, we pursued the analysis with the 67 SOOI categories.

Using Rosvall and Bergstrom's (2008) algorithm with 2006 data, we obtained findings similar to those of these authors on August 11, 2008. Like the original authors using 6,128 journals in 2004, we found 88 clusters using 7,611 journals in 2006.

Lambiotte, one of the coauthors of Blondel et al. (2008), was so kind as to input the data into the unfolding algorithm and found the following results: 114 communities with a modularity value of 0.527708 and 14 communities with a modularity value of 0.60345. We use the 114 communities for the purposes of this comparison. These categories refer to 7,607 (= 7611 − 4) journals because four of the journals in the file were isolates.

The number of journals per category is log-normally distributed in each of the four classifications. In other words, they all have a relatively small number of categories with a large number of journals and many categories with only a few journals. However, as shown in Figure 1, the classifications based on the random walk and unfolding algorithms are more skewed than the content-based classifications.



Whereas the top-10 categories on the basis of a random walk comprise 57% of the journals (50% for unfolding), they cover only 15% in the ISI decomposition and 31% for the SOOI classification. In the case of skewed distributions, the characteristic number of journals per category can best be expressed by the median: the median is below 30 in the random walk or unfolding classifications, compared with 42 journals for the ISI classification and 141 for the SOOI classification (Table 1).


As presented in the last rows of Table 1, the total numbers of citations in the aggregated matrices based on the ISI or SOOI classifications are much higher because the same citation can be attributed to two or three categories. Thus, whereas random walk and unfolding lead to matrices with most citations within categories (on the diagonal), matrices based on ISI and SOOI classifications lead to matrices with most citations between categories (off-diagonal).

Finally, to measure how similar the categories in the four decompositions are to each other, we computed the cosine similarities in the citation patterns between each pair of citing categories in the four aggregated category-category matrices (Salton & McGill, 1983; Ahlgren, Jarneving, & Rousseau, 2003).

We find again that all the distributions are highly skewed and that the random walk and unfolding algorithms exhibit a much lower median similarity value among categories. The lower medians indicate that the algorithmic decompositions produce a much “cleaner” cut between categories than the content-based classifications.
In conclusion, the analysis of the statistical properties of the different classifications teaches us that the random walk and the unfolding algorithms produce much more skewed distributions in terms of the number of journals per category, but these constructs are more specific than the content-based classification of the ISI and SOOI. The content-based sets are less divided because the boundaries among them are blurred by the multiple assignments.

In summary, although the correspondences among the main categories are sometimes as low as 50% of the journals, most of the mismatched journals appear to fall in areas within the close vicinity of the main categories. In other words, it seems that the various decompositions are roughly consistent but imprecise.

Maps of science for each decomposition were generated from the aggregated category-category citation matrices using the cosine as similarity measure.

The similarity matrices were visualized with Pajek (Batagelj & Mrvar, 1998) using Kamada and Kawai's (1989) algorithm.

The threshold value of similarity for edge visualization is pragmatically set at cosine > 0.01 for the algorithmic decompositions and cosine > 0.2 for the content-based decompositions to enhance the readability of the maps without affecting the representation of the structures in the data.

For the ISI decomposition, the 220 categories (Figure 3) were clustered into 18 macro-categories (Figure 4) obtained from the factor analysis (cf. Leydesdorff and Rafols, 2009).


The map of the SOOI classification was constructed with all is 67 subfields (Figure 5).


Taking advantage of the concentration of journals in a few categories, in the case of random walk and unfolding only the top 30 and 35 categories were used, respectively.


Indeed, the four maps correspond in displaying two main poles: a very large pole in the biomedical sciences and a second pole in the physical sciences and engineering. These two poles are connected via three bridging areas: chemistry, a geosciences-environment-ecology group, and the computer sciences. The social sciences are somewhat detached, linked via the behavioral sciences/neuroscience to the biomedical pole, and via the computer sciences and mathematics to the physics/engineering pole.

As noted above, although categories of different decompositions do not always match with one another, most “misplaced” journals are assigned into closely neighbouring categories. Therefore, the error in terms of categories is not large and is also unsystematic. The noise-to-signal ratio becomes much smaller when aggregated over the relations among categories.

As a second important observation that can be made on the basis of these maps, we wish to point to the differences in category density between the content-based and the algorithm-based maps.

In summary, we were surprised to find that the different science maps are similar except that they differ in the density of categories within groups.

The content-based classifications achieve a more balanced coverage of the disciplines at the expense of distinguishing categories that may be highly similar in terms of journals.

The first finding is that the algorithmic decompositions have very skewed and clean-cut distributions, with large clusters in a few scientific areas, whereas indexers maintain more even and overlapping distributions in the content-based classifications.

Second, the different classifications show a limited degree of agreement in terms of matching categories. In spite of this lack of agreement, however, the science maps obtained are surprisingly similar; this robustness is due to the fact that although categories do not match precisely, their relative positions in the network among the other categories is based on distributions that match sufficiently to produce corresponding maps at the aggregated level.

2015年4月6日 星期一

Chen, C.-M. (2008), Classification of scientific networks using aggregated journal-journal citation relations in the Journal Citation Reports. Journal of the American Society for Information Science and Technology, 59(14), 2296–2304. doi: 10.1002/asi.20935

Chen, C.-M. (2008), Classification of scientific networks using aggregated journal-journal citation relations in the Journal Citation Reports. Journal of the American Society for Information Science and Technology, 59(14), 2296–2304. doi: 10.1002/asi.20935

本研究利用親似傳導法(affinity propagation method, Frey & Dueck, 2007),以彙整的期刊對期刊引用關係(aggregated journal-journal citation relation),對期刊間由相似的引用樣式(citation patterns)形成的科學網路進行分類。過去已有許多以期刊對期刊引用資料進行分析的研究,例如Pudovkin and Garfield (2002) 根據引用資料,發展關係係數(relatedness factor)來發現意義相關的期刊(semantically related journals);Doreian and Fararo (1985)發現網路上結構對等(structure equivalence)的期刊;Leydesdorff and Cozzens (1993)利用主成分分析(principal component analysis)取得科學網路的特徵向量(eigenvectors)。本研究所使用的引用資料包括2001年的SCI(共使用1905種期刊、426065篇文章以及13798138個引用資料)以及2005年的SSCI(共使用1578種期刊、66051篇文章以及2437389個引用資料)。本研究所使用的親似傳導法利用s(i,j)= −dij測量期刊j可以做為期刊i所在類別代表期刊的適合性,而dij的計算為

csij則是期刊間的引用樣式(citation pattern)的相似性:


親似傳導法反覆計算期刊間的兩種數值估算期刊間的代表性,r(i, j)反應期刊j能否代表期刊i的適合程度,

a(i, j)則反應期刊i是否應選擇期刊j作為代表的適合程度,


對期刊i來說,最大的a(i, j) + r(i, j)便指明哪一個期刊j可以代表它。

根據分類的結果,一個分類的專指性(specificity)可以從所有的成員期刊到此分類的代表期刊的平均距離來表示,愈小的平均距離表示這個分類具有愈高的專指性。成員之間的相關性(relatedness of category members)則以所有的期刊之間的平均距離來表示,愈小表示成員間彼此愈靠近。
本研究對SSCI期刊的分類結果共分為23個分類,每一個分類大致符合SSCI的主題分類,然而分類裡所有成員的平均距離比SSCI相對應的分類還要小。

Traditional classification methods (Glänzel & Schubert, 2003) are based on subjective analysis, whose output could vary from one person to another. In other words, these methods are more artistic than scientific.

On the other hand, a quantitative approach to classification is usually constructed based on a set of simple rules, which offers robust classification schemes that do not rely on human interference.

The aggregated journal-journal (J-J) citation data in JCR contain extensive information about interjournal citations, which could provide an understanding of the interaction among various scientific disciplines.

Based on JCR citation data, Pudovkin and Garfield (2002) have used an intuitive criterion (relatedness factor) for finding semantically related journals.

To avoid subjective analysis, various quantitative methods have been proposed to construct a robust classification system of scientific journals using JCR citation information.

A variety of techniques for analyzing J-J citation relationships have been reported in the literature to cluster scientific journals (Doreian & Fararo, 1985; Leydesdorff, 1986; Tijssen, De Leeuw, & Van Raan, 1987).

For example, by applying the notion of structure equivalence to analyze a small set of journals, Doreian and Fararo (1985) have delineated a set of blocks, which contain journals. These blocks have a very close correspondence to a categorization of the journals based on their aims and objectives.

More recently Leydesdorff and Cozzens (1993) have developed an optimization procedure that stabilizes approximated eigenvectors of the scientific network from principal component analysis as representations of clusters. This principal component analysis has been further extended to rotated component analysis (Leydesdorff, 2006; Leydesdorff & Cozzens, 1993), which enables one to focus on specific subsets with internal coherence.

An alternative method of cocitation clustering has been investigated in constructing a World Atlas of Sciences for ISI (Garfield, Malin, & Small, 1975; Leydesdorff, 1987; Small, 1999).

In this article, I propose a quantitative approach to classify the scientific network in terms of aggregated J-J citation relations of JCR using the affinity propagation method (Frey & Dueck, 2007).

The method used by ISI in establishing journal categories for JCR is a heuristic approach, in which the journal categories have been manually developed initially. The assignment of journals was based upon a visual examination of all relevant citation data.

As the number of journals in a category grew, subdivisions of the category were then established subjectively.

Although this is a useful approach, a more robust, convenient, and automatic classification scheme is desired.

The citation data analyzed include the SCI of 2001 and the SSCI of 2005, which are directly computed from the extraction of the CD version of the ISI database.

There are 2,195 journals of impact factor greater than 1 in the 2001 SCI. After removing 290 journals that did not publish any articles in 2001, there are 1,905 journals left in our data set, which contains 426,065 articles and 13,798,138 citations.

For the 2005 SSCI, there are 1,583 journals in the database, of which 1,578 journals have nonzero contents. The SSCI database contains 66,051 articles and 2,437,389 citations.

In principle, the dissimilarity between two journals can be visualized by the differences in their citation patterns. In other words, the citation pattern of each journal is represented by a normalized citation vector, and these vectors form a rescaled citation matrix. The dissimilarity (or similarity) in citation between two journals is related to the scalar product of their citation vectors.

For mapping or visualization, coefficients of similarity are converted into distances such that closely related journals are short distances apart and remotely related journals are long distances apart.

The affinity propagation method takes as input a collection of similarities between journals, where the similarity s(i, j) measures how well journal j is suited to be the representative of a journal category for journal i. Since the goal is to minimize squared error, we set s(i, j) = −dij.

There are two types of messages exchanged between journals, including the responsibility r(i, j), which is sent from journal i to candidate representative journal (RJ) j, and the availability a(i, j), which is sent from candidate representative journal j to journal i. Here the responsibility reflects the accumulated evidence for how well-suited journal j is to serve as the representative for journal i, and the availability shows the accumulated evidence for how appropriate it would be for journal i to choose journal j as its representative.

Taking into account other potential representative journals for journal i, the responsibility is computed iteratively as

where the initial value of a(i, j) is set to zero in the first iteration. Similarly, taking into account the support from other journals that journal j should be a representative, the availability is updated by gathering evidence from journals as to whether each candidate representative would make a good representative journal:

To reflect accumulated evidence that journal j is a representative based on the positive responsibilities sent to candidate representative j from other journals, the self-availability is updated as

During the process of affinity propagation, the sum of availability and responsibility can be used to identify the representative journal of emerging journal categories. In other words, for any journal i, the value of j that maximizes a(i, j) + r(i, j) identifies that journal j is its representative.

In our classifications, the level of specificity of a category can be found by looking at its value of DRJ (the average distance of members of a category to its representative journal), and relatedness of category members is implied by the value of DJ-J (the average J-J distance within a category).

To demonstrate the applicability of the affinity propagation method in clustering a complete data set of journals, we first apply it to cluster journals in the 2005 SSCI database.

Here the cutoff parameter t is set to 0.0001, implying that the maximal value of DJ-J (DJ-Jmax) is 100. This choice of t is quite reasonable since the probability distribution (PD), or normalized histogram (bin size is 1), of DJ-J in the unclustered SSCI journal database is mostly between 0 and 30, as shown in Figure 1.



With a choice of DJ-Jmax = 100, the distance between unrelated journals is much larger than that between related journals. In other words, for any journal category, unrelated journals will not be located in the vicinity of its members (each journal is considered as a point in a high-dimensional space). Thus only correlated journals will be grouped together by the affinity propagation method.

However, if DJ-Jmax is too close to 30, the positions of unrelated journals are not well separated and the distortion to the journal positions due to the introduction of the cutoff would affect the clustering of journals.

For the predicted SSCI classification, only those J-J distances within the same category are considered in calculating its PD of DJ-J.

In Figure 1, there are two peaks observed from the statistical curves of PD in DJ-J, where the first peak shows the relatedness between journals within the database (or categories), while the second peak at DJ-J = 100 indicates the irrelevance between journals within the database (or categories).

For the predicted SSCI classification, clearly its first peak in the PD of DJ-J is much more prominent and the peak width is much more narrow than that of the unclustered SSCI database.

On the other hand, its second peak of irrelevance is much smaller than that of the unclustered database.

The probability distribution of the first peak is found to decrease exponentially with DJ-J, i.e., P = P0 exp[−(DJ-J − d0)/ Δ], where P0 is the peak value, d0 is the peak position, and Δ is the decay width. By fitting the statistical data, we find that d0 = 4 and Δ = 9.08 for the unclustered SSCI curve, while d0 = 2 and Δ = 1.72 for the clustered SSCI curve.

The entire journal set of SSCI is decomposed into 23 journal categories.

The relatedness of journals within a category can be seen as the average value of DJ-J within the category, and the specificity of a category is related to the average distance of category members to its RJ.

For any category, a smaller value of DRJ implies a higher level of specificity, and a smaller value of DJ-J implies that journals within a category are more closely related to each other.

In general most categories in our classification scheme have a corresponding category in the ISI classification scheme, and their value of DJ-J seems to be smaller than that of their counterpart in the ISI classification scheme.

When a larger value of the cutoff parameter is used, the maximal distance of DJ-J becomes smaller. ... Since the high-dimensional J-J distance space is now approximated by a high-dimensional sphere of smaller radius, the resolution in clustering journals is higher in this case. Thus the SCI database is expected to be decomposed into more clusters for t = 10−3, compared to the case of t = 10−4. ... Therefore, from comparing clustering results with different values of the cutoff parameter, the relationship among various disciplines can be revealed.

Our results demonstrate that the affinity propagation method can provide a reasonable classification scheme for either a complete database or an incomplete database. This method does not need the number of categories or their size as an input.

Distance between journals is calculated from the similarity of their annual citation patterns with a cutoff parameter to restrain the maximal distance.

Different values of the cutoff parameter lead to different levels of resolution in the classification of journal network. A more coarse-grained classification is obtained when a smaller value of the cutoff parameter (or a larger maximal J-J distance) is used.

We note that, unlike the ISI classification scheme, which allows overlap in the content of journal categories by subjective decisions, each journal uniquely belongs to a category in our classification scheme.