顯示具有 citation analysis 標籤的文章。 顯示所有文章
顯示具有 citation analysis 標籤的文章。 顯示所有文章

2016年7月11日 星期一

Wang, Q., & Waltman, L. (2016). Large-scale analysis of the accuracy of the journal classification systems of Web of Science and Scopus. Journal of Informetrics, 10(2), 347-364.

Wang, Q., & Waltman, L. (2016). Large-scale analysis of the accuracy of the journal classification systems of Web of Science and Scopus. Journal of Informetrics10(2), 347-364.

本研究以引用資料比較Web of Science和Scopus兩個資料庫提供的期刊分類系統(journal classification systems)的正確性。分類系統能應用於各種問題;例如,它可以被用來標定研究區域(Glänzel & Schubert, 2003; Waltman & Van Eck, 2012),評估和比較研究對各領域的影響(Leydesdorff and Bornmann, 2015; Van Eck, Waltman, Van Raan, Klautz, & Peul, 2013),以及跨學科的研究(Porter & Rafols, 2009; Porter, Roessner, & Heberger, 2008)。除了Web of Science和Scopus之外,期刊分類系統還有Science-Metrix、NSF(National Science Foundation)分類系統、the UCSD (University of California, San Diego)分類系統以及ANZSRC (Australian and New Zealand Standard Research Classification),另外Glänzel and Schubert (2003)也提出一個包括期刊和論文的階層式分類系統,以演算法建構的期刊分類方法也有Bassecoulard and Zitt (1999)、Chen (2008)以及 Rafols and Leydesdorff (2009)等,Waltman and Van Eck (2012)的演算法則是以期刊裡出版的論文分類為主。

根據Waltman (2015, Section 3)的文獻分析,在比較Web of Science和Scopus時,主要針對資料庫的覆蓋情形(the coverage of the databases),例如LópezIllescas, De Moya-Anegón, & Moed (2008)、 Meho & Rogers (2008)、 Mongeon & Paul-Hus (2016)、 Norris & Oppenheim (2007),或是資料庫用來評估研究產生與影響的準確性,如Archambault, Campbell, Gingras, & Larivière (2009)、Bar-Ilan, Levene, & Lin (2007)、Meho & Rogers (2008)、Meho & Sugimoto (2009),並未有研究比較與分析它們的分類系統的準確性。

過去Pudovkin and Garfield (2002)曾說明WoS首先利用人工的經驗法則將期刊分配到各類別,之後使用根據引用資料的Hayne-Coulson演算法對新的期刊進行分類。除此之外,Katz and Hicks (1995)、Leydesdorff (2007)、Leydesdorff and Rafols (2009)等研究也曾指出WoS的分類系統是綜合了引用模式、期刊題名與專家意見。但Scopus則未曾有文獻提到其分類系統的建構方式。事實上,在WoS上有兩個分類系統,一個為具有250個類別的類別系統(a system of categories),另一個則是包含約150個研究領域的研究領域系統(a system of research areas),此外,另一個分類系統僅包含科學與社會科學,稱為ESI (Essential Science Indicators)。本研究的分析對象是WoS上的類別系統。

Scopus的期刊分類系統則名為ASJC ( All Science Journal Classification),分為兩個層級,下層有304個類別,上層則分為27個類別。

既然期刊分類系統相當有用,因此有許多研究提出對WoS及Scopus的分類系統進行改善的方法,例如Glänzel等人研究各種方法來驗證與改善WoS的分類系統 (Janssens, Zhang, De Moor, & Glänzel, 2009; Thijs, Zhang, & Glänzel, 2015; Zhang, Janssens, Liang, & Glänzel, 2010),López-Illescas, Noyons, Visser, De Moya-Anegón, & Moed (2009) 則是對利用WoS分類系統進行的領域劃分,提出改進的方法。在Scopus分類系統的改進方面,則有SCImago團隊(Gómez-Núnez, ˜ Vargas-Quesada, De Moya-Anegón, & Glänzel, 2011, Gómez-Núnez, ˜ Batagelj, Vargas-Quesada, De Moya-Anegón, & Chinchilla-Rodríguez, 2014, Gómez-Núnez, ˜ Vargas-Quesada, & De Moya-Anegón, 2016)。

用來評估期刊分類系統準確性的方法可分為以專家為基礎的方法與書目計量方法(bibliometric approach)。以專家為基礎的方法在遭遇大量資料時有很大的困難,沒有專家有足夠知識來評估所有科學學科上的期刊分類,因此需要相當多的專家加入。書目計量方法可再分為以文本為基礎與以引用為基礎兩種方法,分別以在同一類別下期刊論文的文本相似性與引用模式的相似性大小做為衡量期刊是否應在同一類別的標準。本研究採用的是直接引用(direct citation)關係,先前Klavans and Boyack (2015)曾經利用直接引用關係建構論文的分類系統演算法,他們的結論認為直接引用較書目耦合(bibliographic coupling)或共被引(co-citation)等間接地引用關係更加準確。

綜合以上所述,本研究提出方法的原理可歸納為任一期刊引用該期刊所屬類別下的期刊或被這些期刊所引用的頻次必然較其他類別的期刊之頻次來得高。根據這樣基本原則,本研究制訂兩個檢驗期刊是否被指定於適當類別的標準:
標準1:某一期刊與它所屬類別的其他期刊之間若是只有相當少的引用關係,則這個期刊的分類可能有問題。
標準2:如果某一期刊與其他類別的期刊之間有相當多的引用關係,則這個期刊可能被分到不正確的類別下。

本研究以2010到2014年的WoS及Scopus資料庫上的所有期刊作為分析資料,相關的統計數據如表1所示,比較兩者所收錄的資料,Scopus資料庫除了比WoS更多的期刊種類與類別以外,每一種期刊被指定的類別數目也通常比較多,WoS上的期刊平均被指定到約1.6個類別,但是Scopus則為2.1。


根據第一項標準,WoS及Scopus兩個資料庫上都有許多期刊被指定到不適合的類別上,並且Scopus尤為嚴重,如表3所示。在各個至少有10種期刊的類別當中,選擇至少有一半的期刊符合標準一的類別,這樣的類別,WoS共有17個,Scopus則有高達76個,在兩個資料庫上都出現的類別包括建築學(ARCHITECTURE)、生物物理學(BIOPHYSICS)以及 醫學實驗室技術(MEDICAL LABORATORY TECHNOLOGY)。



兩個資料庫的情況在標準二上都有還不錯的結果(表6)。


同時符合標準一與標準二的期刊,一方面與本身被指定的類別只有較弱的連結,另一方面則與未被指定的類別有較強的連結,分析同時符合標準一與標準二的期刊可以發現,一種可能是在這些期刊上的發表已經和它們的題名和範圍宣告(scope statement)有所差異,另一種可能則是在分類時僅依賴它們的題名。

根據以上的實驗結果,可以歸納以下的幾點結論:
1. 在標準一,WoS的表現比Scopus還要好,因此可以說,Scopus上的期刊通常與它們被指定的類別只有較弱的連結。
2. 在標準二,兩個資料庫的表現都相當好,也就是如果某一期刊與某一個類別的連結較強的話,WoS及Scopus通常會將它指定到這個類別。
3. 整合兩個標準,WoS比Scopus的表現通常要好上許多。

除了上述的結論外,本研究還指出Scopus有些類別有容易混淆的名稱,例如有兩個類別分別命名為LINGUISTICS & LANGUAGE和LANGUAGE & LINGUISTICS,另兩個則為INFORMATION SYSTEMS & MANAGEMENT與MANAGEMENT INFORMATION SYSTEMS。

而且兩個分類系統都缺乏透明性,本研究的作者沒有發現建構與更新分類系統的適當文件。


To examine and compare the accuracy of journal classification systems, we define two criteria on the basis of direct citation relations between journals and categories. We use Criterion I to select journals that have weak connections with their assigned categories, and we use Criterion II to identify journals that are not assigned to categories with which they have strong connections. If a journal satisfies either of the two criteria, we conclude that its assignment to categories may be questionable.

Accordingly, we identify all journals with questionable classifications in Web of Science and Scopus. Furthermore, we perform a more in-depth analysis for the field of Library and Information Science to assess whether our proposed criteria are appropriate and whether they yield meaningful results.

It turns out that according to our citation-based criteria Web of Science performs significantly better than Scopus in terms of the accuracy of its journal classification system.

Classifying journals into research areas is an essential subject for bibliometric studies.

A classification system can assist with various problems; for instance, it can be used to demarcate research areas (e.g., Glänzel & Schubert, 2003; Waltman & Van Eck, 2012), to evaluate and compare the impact of research across scientific fields (e.g., Leydesdorff and Bornmann, 2015; Van Eck, Waltman, Van Raan, Klautz, & Peul, 2013), and to study the interdisciplinarity of research (e.g., Porter & Rafols, 2009; Porter, Roessner, & Heberger, 2008).

Besides the WoS and Scopus classification systems, there are various other multidisciplinary classification systems, for instance the system of Science-Metrix,the system of the National Science Foundation (NSF) in the US,the UCSD classification system, and the system of the Australian and New Zealand Standard Research Classification (ANZSRC).

Science-Metrix assigns “individual journals to single, mutually exclusive categories via a hybrid approach combining algorithmic methods and expert judgment” (Archambault, Beauchesne, & Caruso, 2011, p. 66). The Science-Metrix system includes 176 categories.

The NSF system also offers a mutually exclusive classification of journals, but it is more aggregated, consisting of only 125 categories (Boyack & Klavans, 2014). The system is used in the Science & Engineering Indicators of the NSF.

A more detailed classification system is the so-called University of California, San Diego (UCSD) classification system. This system, which includes more than 500 categories, has been constructed in a largely algorithmic way. The construction of the UCSD classification system is discussed by Börner et al. (2012).

The ANZSRC’s Field of Research (FoR) classification system has a three-level hierarchical structure. Journals are classified at the top level and at the intermediate level. Journals can have multiple classifications.

Furthermore, Glänzel and Schubert (2003) designed a two-level hierarchical classification system, which can be applied at the levels of both journals and publications. They adopted a top-bottom strategy; specifically, they first defined categories on the basis of the experience of bibliometric studies and external experts. They then assigned journals and individual publications to the categories. This classification system has for instance been used for measuring interdisciplinarity. In their analysis of interdisciplinarity, Wang, Thijs, & Glänzel (2015) explain that instead of the WoS subject categories they use the more aggregated classification system developed by Glänzel and Schubert (2003).

Algorithmic approaches to construct classification systems at the level of journals have been studied by for instance Bassecoulard and Zitt (1999), Chen (2008), and Rafols and Leydesdorff (2009).

A more recent development is the algorithmic construction of classification systems at the level of individual publications rather than journals. Waltman and Van Eck (2012) developed a methodology for algorithmically constructing classification systems at the level of individual publications on the basis of citation relations between publications. Their approach has for instance been used in the calculation of field-normalized citation impact indicators (Ruiz-Castillo & Waltman, 2015).

According to a recent literature review (Waltman, 2015, Section 3), previous studies comparing WoS and Scopus are mainly focused on two aspects. One is the coverage of the databases (e.g., LópezIllescas, De Moya-Anegón, & Moed, 2008; Meho & Rogers, 2008; Mongeon & Paul-Hus, 2016; Norris & Oppenheim, 2007) and the other is the accuracy of the databases when used to assess research output and impact at different levels, ranging from individual researchers to departments, institutes, and countries (e.g., Archambault, Campbell, Gingras, & Larivière, 2009; Bar-Ilan, Levene, & Lin, 2007; Meho & Rogers, 2008; Meho & Sugimoto, 2009). However, no study has systematically compared WoS and Scopus in terms of the accuracy of their journal classification systems.

In the case of WoS, Pudovkin and Garfield (2002) have offered a brief description of the way in which categories are constructed. According to Pudovkin and Garfield, when WoS was established, a heuristic and manual method was adopted to assign journals to categories, and after this, the so-called Hayne-Coulson algorithm was used to assign new journals. This algorithm is based on a combination of cited and citing data, but it has never been published.

Besides this, Katz and Hicks (1995), Leydesdorff (2007), and Leydesdorff and Rafols (2009) have indicated that the WoS classification system is based on a comprehensive consideration of citation patterns, titles of journals, and expert opinion.

In the case of Scopus, there seems to be no information at all on the construction of its classification system.

It should be mentioned that in the most recent versions of WoS two classification systems are available, namely a system of categories and a system of research areas.

The system of categories is more detailed. This system, which is the traditional classification system of WoS and the system on which we focus our attention in this paper, consists of around 250 categories and covers the sciences, social sciences, and arts and humanities.

The system of research areas, which has become available in WoS more recently, is less detailed and comprises around 150 areas.

Besides these two systems, Thomson Reuters also has a classification system for its Essential Science Indicators. This system consists of 22 subject areas in the sciences and social sciences. It does not cover the arts and humanities.

The Scopus journal classification system is called the All Science Journal Classification (ASJC). It consists of two levels. The bottom level has 304 categories, which is somewhat more than the about 250 categories in the WoS classification system. The top level includes 27 categories.

The accuracy of a classification system can seriously influence bibliometric studies. For instance, Leydesdorff and Bornmann (2015) investigated the use of the WoS categories for calculating field-normalized citation impact indicators. They focused specifically on two research areas, namely Library and Information Science and Science and Technology Studies. Their conclusion is that “normalizations using (the WoS) categories might seriously harm the quality of the evaluation”.

A similar conclusion was reached by Van Eck et al. (2013) in a study of the use of the WoS categories for calculating field-normalized citation impact indicators in medical research areas.

Glänzel and colleagues have studied several approaches to validate and improve WoS-based classification systems (Janssens, Zhang, De Moor, & Glänzel, 2009; Thijs, Zhang, & Glänzel, 2015; Zhang, Janssens, Liang, & Glänzel, 2010). They have also proposed an improved way of handling publications in multidisciplinary journals (Glänzel, Schubert, & Czerwon, 1999;Glänzel, Schubert, Schoepflin, & Czerwon, 1999).

Related to this, López-Illescas, Noyons, Visser, De Moya-Anegón, & Moed (2009) have studied an approach to improve the field delineation provided by categories in the WoS classification system.

The SCImago research group has made a number of attempts to improve the Scopus classification system (Gómez-Núnez, ˜ Vargas-Quesada, De Moya-Anegón, & Glänzel, 2011, Gómez-Núnez, ˜ Batagelj, Vargas-Quesada, De Moya-Anegón, & Chinchilla-Rodríguez, 2014, Gómez-Núnez, ˜ Vargas-Quesada, & De Moya-Anegón, 2016).

Two types of approaches can be distinguished for assessing the accuracy of journal classification systems. One is the expert-based approach and the other is the bibliometric approach.

Applying the expert-based approach at a large scale is challenging. No expert has sufficient knowledge to assess the classification of journals in all scientific disciplines, so a large number of experts would need to be involved.

In the case of the bibliometric approach, a further distinction can be made between text-based and citation-based approaches.

Text-based approaches could for instance assess whether the textual similarity of publications in journals assigned to the same category is higher than the textual similarity of publications in journals assigned to different categories.

Instead, we take a citation-based approach to assess the accuracy of journal classification systems.

In this paper, we use direct citation relations. This is because “a co-citation or bibliographic coupling relation requires two direct citation relations” (Waltman & Van Eck, 2012, p. 2380), which means that bibliographic coupling and co-citation relations are more indirect signals of the relatedness of journals than direct citation relations.

The use of direct citation relations is also supported by Klavans and Boyack (2015), who study the algorithmic construction of classification systems at the level of individual publications. They conclude that the use of direct citation relations yields more accurate results than the use of bibliographic coupling or co-citation relations.

Thus, the rationale of our approach can be summarized as follows: A journal should cite or be cited by journals within its own category with a high frequency in comparison with journals outside its category.

Based on this basic principle, we define two criteria to identify journals with questionable classifications. One criterion is that if a journal has only a very small number of citation relations with other journals within its own category, then we believe the classification of the journal to be questionable. The other criterion is that if a journal has many citation relations with journals in a category to which the journal itself does not belong, then it seems likely that the journal incorrectly has not been assigned to this category.

We retrieved from the WoS and Scopus databases all journals that have publications between 2010 and 2014....The choice of a five-year time window is a trade-off between on the one hand the stability of journal classification systems and on the other hand the accuracy of our approach based on direct citation relations.

As can be seen in Table 1, the number of Scopus journals included in the analysis is almost twice as large as the number of WoS journals, and Scopus also includes 80 more categories than WoS. Furthermore, although both databases often assign journals to multiple categories, we found that Scopus tends to assign journals to more categories than WoS. WoS assigns journals to at most six categories, whereas in Scopus there turns out to be a journal that is assigned to 27 categories. Additionally, we found that the average number of categories to which journals belong equals 1.6 in WoS and 2.1 in Scopus. This shows that on average journals have significantly more category assignments in Scopus than in WoS.

As can be seen, almost 60% of all journals in WoS belong to only one category, whereas in Scopus more than 60% of all journals are assigned to two or more categories.


WoS has 1390 journals with ti < 100, accounting for 11% of the total number of WoS journals, whereas Scopus has 5808 journals with ti < 100, which is 24% of the total.3 Hence, Scopus has more journals with ti < 100 than WoS not only in an absolute sense but also from a relative point of view.

Taking a further look at Scopus journals with ti < 100, it turns out that they can be roughly divided into three groups. One group consists of arts and humanities journals, another group consists of newly included journals, and a third group consists of non-English language journals.

Table 2 provides some basic statistics on the assignment of journals to categories in WoS and Scopus when journals with ti < 100 and assignments of journals to multidisciplinary categories are excluded. The table shows the number of journals that belong to at least one non-multidisciplinary category and the number of assignments of journals to non-multidisciplinary categories. As can be seen in the table, in the case of Scopus the constraints that we have introduced cause a much larger decrease in the number of journals and the number of journal-category assignments than in the case of WoS.

Table 3 reports for both WoS and Scopus and for three values of the threshold˛the number of journals and the number of journal-category assignments that satisfy Criterion I.

As can be seen, both databases have assigned a significant number of journals to categories that according to Criterion I seem to be inappropriate.

Moreover, no matter which threshold is considered, Scopus performs substantially worse than WoS, not only in the absolute number of journals and journal-category assignments satisfying Criterion I but, more importantly, also in the percentage of journals and journal-category assignments satisfying the criterion.

Next, we identify WoS and Scopus categories with a high percentage of journals satisfying Criterion I. The identified categories may be seen as the most problematic categories in the two databases, because many of the journals belonging to these categories are only weakly connected to each other in terms of citations.

We select categories that include at least 10 journals with ti ≥ 100 and that, for α˛= 0.1, have at least 50% of their journals satisfying Criterion I. The results for WoS and Scopus are reported in Tables 4 and 5, respectively. In the case of WoS 17 categories have been identified, whereas in the case of Scopus 76 categories have been identified, so more than four times as many as in the case of WoS.

There are three categories that have been identified in the case of both databases: ARCHITECTURE, BIOPHYSICS, and MEDICAL LABORATORY TECHNOLOGY.

Table 6 presents for both WoS and Scopus and for five values of the threshold ˇ the number of journals that satisfy Criterion II.



A journal satisfies both Criterion I and Criterion II if on the one hand it has weak connections, in terms of citations, with its assigned categories while on the other hand it has a strong connection with a category to which it is not assigned. More precisely, our focus is on journals for which the current category assignments all satisfy Criterion I, while there is an alternative category assignment that satisfies Criterion II.

Based on the three journals discussed above, we conclude that journals satisfying the combined Criteria I and II can be classified into at least two types. One type refers to journals for which there is a discrepancy between on the one hand their title and their scope statement and on the other hand what they have actually published.  ... The second type refers to journals that seem to have been assigned to a category based only on their title.

First, WoS performs much better than Scopus according to Criterion I. Using the parameter values ˛= 0.05 and ˛= 0.1, the percentage of journals and journal-category assignments satisfying Criterion I is more than two times higher for Scopus than for WoS. Hence, in Scopus journals are assigned to categories with which they are only weakly connected much more frequently than in WoS.

Second, based on Criterion II, WoS and Scopus both perform reasonably well, with WoS having a somewhat better performance than Scopus. For all parameter values that were considered, less than 5% of all journals in WoS and Scopus satisfy Criterion II. In other words, if a journal is strongly connected to a category, WoS and Scopus typically assign the journal to that category.

Third, WoS also presents a significantly better result than Scopus based on the combined Criteria I and II. In WoS there is only one journal satisfying the combined criteria, whereas in Scopus there are 32.

First, Scopus sometimes has confusing category labels. In particular, Scopus sometimes has two categories with very similar labels. Examples are the categories LINGUISTICS & LANGUAGE and LANGUAGE & LINGUISTICS and the categories INFORMATION SYSTEMS & MANAGEMENT and MANAGEMENT INFORMATION SYSTEMS.

Second, lack of transparency is a weakness of both the WoS and the Scopus classification system. We did not find proper documentation of the methods used to construct and update the WoS and Scopus classification systems.

For instance, in the case of a small category, it may be hardly possible for a journal to have a reasonably high relatedness with the category. Therefore it can be expected that many journals belonging to the category will satisfy Criterion I. This may be caused not so much by the misclassification of these journals but more by the small size of the category. On the other hand, in the case of a large category, there may be other problems. A large category may for instance be of a heterogeneous nature and may cover multiple fields that are hardly connected to each other.

2016年7月7日 星期四

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

本研究利用引用文獻出現文句內容的相似性來測量作者的主題相關性(topical relatedness)。傳統的作者共被引分析(Author co-citation Analysis, ACA)做法是利用參考文獻裡被引用作者的共被引頻率(White and Griffith, 1981),然後利用Pearson相關係數 (Pearson correlation coefficient)或是 Salton提出的餘弦相似性測量作者的相似性,在書目計量學研究裡已經廣泛運用於確認與追蹤學科的知識結構(the intellectual structure of an academic discipline) (He & Hui, 2002)。然而這種做法並未考慮引用的內容,Jeong, Song, & Ding, (2014)與 Zhao & Strotmann (2014)則利用全文裡提到的作者並將有關的內容加入ACA的計算。

本研究認為累積被引作者出現的文句能夠代表作者的研究領域,因此利用JASIST的全文資料,剖析HTML,取出論文的後設資料(題名、作者姓名、出版年、DOI與摘要)、引用資訊(引用文句與參考文獻索引)以及參考資訊(作者姓名、出版年、題名與期刊)。在這個研究裡,共使用2003年1月到2015年6月的1910篇論文,合計77,408筆參考文獻。將引用文句與一般文句分開,連結文句內的參考文獻索引與參考資訊,選取100位最多被引用的作者,進行傳統的ACA以及本研究提出的新方法。本研究的新方法利用Mikolov et al., (2013)提出的Word2Vec 模型 (Word2Vec models),根據參考文獻出現的引用文句,找出作者間的相似性。Word2Vec 模型以大量的文本為基礎,利用類神經網路方法( neural network approaches),找出詞語之間的語意關係,將每一個出現於文句的詞語轉換成向量,使得這些向量之間的相似性能夠保持詞語在語意上的關係。本研究將被引用的作者姓名視為是引用文句中出現的詞語,測量作者間在研究主題的相似性與合作關係。

表2是傳統的ACA方法與本研究的方法分別找出的10組最相似的作者,本研究的方法找出10組最相似的作者中有一半是具有合作關係的作者。

另外,將兩種方法產生的作者關係分別繪製成網路圖,節點代表作者,利用PageRank決定的節點大小,節點的遠近由作者間的相似性決定,並且以Blondel, Guillaume, Lambiotte, & Lefebvre (2008)提出的社群偵測(community detection)方法進行分群。圖三與圖四分別是傳統ACA與本研究提出方法的結果。

圖三上可明顯地看到所有的作者分為兩群,依據社群偵測的分群結果,左邊的作者可再分為兩群:最左邊紅色的一群為研究資訊尋求行為(information seeking behavior)的作者,紫色的一群則與資訊檢索(information retrieval)研究有關,右邊綠色的一群則是研究書目計量學(bibliometrics)的作者。介於左右兩大群體的作者分別有兩位:Borgman和Salton。這兩位都是資訊科學領域傳統上會經常引用的作者。



在以Word2Vec方法產生的作者網路上,與資訊檢索有關的作者群組位於左方,包括上方的資訊尋求行為以及下方的文件檢索(document retrieval)兩個群組,書目計量學在圖四上則分為兩個有關的群組,一個主要包含作者分析(author analysis),另一則是期刊引用分析(journal citation analysis)與評鑑指標(evaluation indicator)。與圖三不同的是,圖四上的群組彼此間都有連結,並且圖形上更具體地呈現次學科(sub-disciplines)以及重要的作者。

Unlike other ACA studies, we used citing sentences to reflect topical relatedness of authors.

In  our  research,  we extended  traditional  approaches by  adopting Word2Vec, one  of  deep learning methods, to measure author similarity.

We also conducted in-depth network analysis of author maps.

The results of Word2Vec-based author map revealed more specific sub-disciplines and the important authors in perspective of topical influence than traditional approach does.

Author co-citation Analysis (ACA), which was introduced by White and Griffith (1981), has been widely used in bibliometrics researches to identify and trace the intellectual structure of an academic discipline (He & Hui, 2002). In ACA, traditional approaches relied on the co-citation frequency of cited authors in the reference section.

Thus, one of the main topics in ACA was methodological discussion of what kind of measure is appropriate and relevant for calculation of author similarities (Leydesdorff, 2005; van Eck & Waltman, 2007). Existing approaches based on co-citation frequencies such as Pearson correlation coefficient and Salton’s cosine similarity, however, do not capture the citation content.

Thus, some recent researches used the full-text to obtain the topical relatedness between the cited authors (Jeong, Song, & Ding, 2014; Zhao & Strotmann, 2014). They analyzed the authors mentioned in the full-text and incorporated contents related with cited authors into ACA.

In that sense, cumulated citing sentences of cited authors are able to well represent the cited researches and cited authors’ research areas. In addition, these citing sentences are particularly useful for summarization of a research document.

Figure 1 shows the overall system flow of our approach.


For content analysis, however, we collected full-text research articles of JASIST in HTML format. Through the HTML parsing process, we extracted the metadata (title, author name, year, DOI and abstract), citation information (citing sentence, and reference id), and reference information (author name, year, title, and journal).

To compare our method to traditional ACA, we computed author-pairs in both approaches. In Word2Vec-based method, the full-text data, first, are splitting into sentences. In second step, matching the citing sentences with reference id in reference section, we separated the citing sentences and other general sentences. Then, citing sentences are preprocessed in the following steps: tokenization, POS tagging, lemmatization of the tokenized sentence, and stop word removal.

From these data, we trained Word2Vec model for calculating author similarity and generated author-author similarity matrix. To compare the previous research, traditional author counting approach, we also construct co-citation matrix based on citation counts. Since we preprocessed full-text including all reference information, these matrices considered all cited authors.

To evaluation, we selected top 100 authors which are highly cited in both methodology, and conduct network analysis through visualizing author maps.

The data was gathered from 1,910 full-text articles in the JASIST digital library over 12 years (from January 2003 to June 2015). The 1,910 collected documents have 77,408 references. We extracted elements from the full-text article: 1) citing sentences from the body of the article, 2) the references information, and 3) all cited authors. Table 1 shows the basic statistics of collected data.

Word2Vec models, one of the neural network approaches, are able to carry semantic meanings and turns text into a numerical form that deep-learning nets can understand (Mikolov et al., 2013). Based on a large amount of plain text, Word2Vec trains relationships between words automatically.

Word2Vec spatially encoded a word meaning and the relationship between words, which was originally applied to word clustering or synonym detection (Wolf et al., 2014). We applied Word2Vec into author similarity measure regarding cited author names as a word in plain text.

Since authors’ oeuvre was represented as the citing sentences in research articles, the Word2Vec-based method could consider various topics of the author.

In the proposed approach, however, the author names are also trained as words in a same citing sentence. Therefore, the similarity between two authors in the Word2Vec-based method reflects both topical relatedness and collaborations.

Table 2 shows top 10 pairs by the traditional ACA method (Pearson correlation based similarity) and the Word2Vec based approach respectively. About the half of pairs resulted from the Word2Vec approach are the co-author relationship.

This results imply that the proposed approach enables to detect wider range of author pairs in perspective of topical relatedness and grasp more diverse research fields of information science.

To examine whether there are structural differences in two measures of author similarity, we constructed two author networks with top 100 authors. For network visualization, we used PageRank (Brin & Page, 1998) to determine the node size and also adopted the modularity algorithm (Blondel, Guillaume, Lambiotte, & Lefebvre, 2008) for the community detection.

Figure 3 illustrates roughly two parts that consist of information retrieval and bibliometrics, two major research areas in JASIST. The author group of information retrieval (purple) along with information seeking behavior (red) is located at the left side, and the author group related with bibliometrics is located at the right side.

There are only two authors located between two groups (Borgman and Salton), who are traditionally cited authors in the information science field. Borgman studied various topics including information retrieval and scholarly communication and wrote the important books that had won the best information science book from ASIST. Salton’s works also received a lot of citations for a long time in the field of information science.


The author group related with information retrieval in the left side of the network is split into information seeking behavior (blue) located in the upper side of the network and document retrieval (yellow) located at the bottom side of the network. The group related to bibliometrics is also separated into two parts: (1) a cluster (green) including author analysis and (2) journal citation analysis and evaluation indicator (red).

Unlike Figure 3, the communities in the network are connected to each other. Brin is connected with both document retrieval and citation analysis communities. This may be attributed to the fact that the PageRank, developed by Brin and Page (1998), is used in information retrieval and also studied in network analysis to compute node centrality.

In bibliometrics, PageRank is adopted as one of the centralities in citation networks (Ding, Yan, Frazho, & Caverlee, 2009). Ingwesen, who is located between information retrieval and bibliometrics, studied information retrieval in earlier works, he extended the research area to network analysis such as webometrics.

It implies that the authors linked by citation are topically grouped in the Word2Vec-based author network.

2015年4月15日 星期三

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V. (2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V.(2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

由於一般認為將領域之間的關係表示為圖形,通過考慮這些關係的可能性能夠提供許多資訊,不論對新進人員或專家皆有助於理解與分析,因此對這方面方法與工具的需求逐漸提高。過去的研究大多以期刊為分析單位,產生所有科學研究領域的科學映射圖。例如Leydesdorff (2004a, 2004b)使用雙重連結成分(biconnected components)的圖形分析演算法,將JCR 2001的科學研究進行分類。Boyack, Klavans, and Börner (2005)則應用了8種不同的期刊相似性測量7121種SCI和SSCI期刊,並採用VxOrd產生科學映射圖。Samoylenko, Chao, Liu, and Chen (2006)建構科學期刊的最小生成樹(minimum spanning trees),他們使用的資料是SCI 1994到2001的資料。本研究提出一個將ISI (Institute of Scientific Information)類別繪製成科學映射圖的方法,這個方法利用根據類別間的共被引資訊建構類別間的連結,以尋徑網路(PathfinderNetwork)縮減不重要的連結,然後以Kamada-Kawai方法決定節點在圖上的布局(layout),最後利用因素分析(factor analysis)進行結構確認。本研究和先前的研究都是針對類別利用共被引資訊呈現科學映射圖。以類別為分析單位在代表上足夠明確,並且比起較小的單位,這種方式對非專家使用者(nonexpert user)較具有資訊且使用者友善。Moya-Anegón et al. (2004)針對西班牙科學研究領域的視覺化,Moya-Anegón et al. (2005)則進一步利用科學映射圖比較英國、法國和西班牙三個國家的科學研究領域。本研究依循Börner, Chen, and Boyack (2003)提出的知識領域映射流程。使用的資料為7585種ISI期刊,ISI的類別共有219個,但扣除多學科科學後(Multidisciplinary Sciences),採用的類別共218個。利用共被引計算期刊相似性的方式為

Cc(ij)為期刊i和期刊j共被引次數,c(i)和c(j)則分別是期刊i和期刊j被引用次數。然後以尋徑網路和Kamada-Kawai方法繪製網路圖,經過尋徑網路處理後,有較多連結的節點具有較重要的地位。而尋徑網路是一種以型態為主的方法,與以群集為主的因素分析彼此間可以互補,因素分析可以識別、界定與定名科學映射圖上呈現的主題區域,而尋徑網路則負責讓使主題區域更加明顯,將類別分組成束,並顯示連接不同顯著類別的路徑,以及總體的型態結構。。最後總計共分析出35個因素,通過陡坡考驗(scree test)則有16個。科學映射圖上的類別可以分為三個群集:醫學與地球科學、基礎與實驗科學以及社會科學。

This study proposes a new methodology that allows for the generation of scientograms of major scientific domains, constructed on the basis of cocitation of Institute of Scientific Information categories, and pruned using PathfinderNetwork, with a layout determined by algorithms of the spring-embedder type (Kamada–Kawai), then corroborated structurally by factor analysis.

We present the complete scientogram of the world for the Year 2002.

This need arises from the general conviction that an image or graphic representation of a domain favors and facilitates its comprehension and analysis, regardless of who is on the receiving end of the depiction and whether a newcomer or an expert.

Science maps can be very useful for navigating around in scientific literature and for the representation of its spatial relations (Garfield, 1986). They are optimal means of representing the spatial distribution of the areas of research while also offering additional information through the possibility of contemplating these relationships (Small & Garfield, 1985).

From a general viewpoint, science maps reflect the relationships between and among disciplines; but the positioning of their tags clues us into semantic connections while also serving as an index to comprehend why certain nodes or fields are connected with others.

Moreover, these large-scale maps of science show which special fields are most productively involved in research—providing a glimpse of changes in the panorama—and which particular individuals, publications, institutions, regions, or countries are the most prominent ones (Garfield, 1994).

It is a tool in that it allows the generation of maps, and a method in that it facilitates the analysis of domains, by showing the structure and relations of the inherent elements represented. In a nutshell, scientography is a holistic tool for expressing the discourse of the scientific community it aspires to represent, reflecting the intellectual consensus of researchers on the basis of their own citations of scientific literature.

In Moya-Anegón et al. (2004), we ventured forth with a historic evolution of scientific maps from their origin to the present, and proposed ISI-JCR category cocitation for the representation of major scientific domains. Its utility was demonstrated by a visualization of the scientific domain of geographical Spain for the Year 2000.

Since then, other works related with the visualization of great scientific domains have appeared; however, all use journals as the unit of analysis, with the exception of a study based on the cocitation of categories (Moya-Anegón et al., 2005), comparatively focusing on three geographic domains (England, France, and Spain).

In contrast, Leydesdorff (2004a, 2004b) classified world science using the graph-analytical algorithm of biconnected components in combination with JCR 2001.

Boyack, Klavans, and Börner (2005) applied eight alternative measures of journal similarity to a dataset of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes, to show the first global map of science using the force-directed graph layout tool VxOrd.

Samoylenko Chao, Liu, and Chen (2006) proposed an approach through the construction of minimum spanning trees of scientific journals, using the Science Citation Index from 1994 to 2001.

In processing and depicting the scientific structure of great domains, we further developed a methodology that follows the flow of knowledge domains and their mapping as proposed by Börner, Chen, and Boyack (2003).

Because ISI assigns each journal to one or more subject categories, to designate a subject matter (i.e., ISI category) for each document, we also downloaded the Journal Citation Report (JCR; Thomson Corporation, 2005a), in both its Science and Social Sciences editions, for 2002.

The downloaded records were exported to a relational database that reflects the structured information of the documents. This new repository contained nearly 1 million (N = 901,493) source documents: articles, biographical items, book reviews, corrections, editorial materials, letters, meeting abstracts, news items, and reviews that had been published in 7,585 ISI journals (N = 5,876 + 1,709). These were classified in a total of 219 categories, altogether citing 25,682,754 published documents.

As informational units, they are, in themselves, sufficiently explicit to be used in the representation of all disciplines that make up science in general. These categories, in combination with the adequate techniques for the reduction of space and the representation of the information to construct scientograms of science or of major scientific domains, prove much more informative and user friendly for quick comprehension and handling by nonexpert users than those obtained by the cocitation of smaller units of cocitation.

For these reasons, we used the 219 categories of the JCR 2002 as units of measure, with the exception of “Multidisciplinary Sciences.” ... The maximum number of categories with which we worked, then, was 218.

In light of our previous experience (Moya-Anegón et al., 2004, 2005), we use cocitation as the similarity measure to quantify the relationship existing between each one of the JCR categories.

Therefore, after a number of trials, we arrived at the conclusion that using tools of Network Analysis, the best visualizations are those obtained through raw data cocitation as the unit of measure. Yet, it also was necessary to reduce the number of coincident cocitations to enhance pruning algorithm yield. Therefore, to those raw data values we added the standardized cocitation value. In this way, we could work with raw data cocitation while also differentiating the similarity values between categories with equal cocitation frequencies. The key was a simple modification of the equation for the standardization of the degree of citation proposed by Salton and Bergmark:




where CM is cocitation measure, Cc is cocitation frequency, c is citation, and i and j are categories.

Over the history of the visualization of scientific information, very different techniques have been used to reduce n-dimensional space. Either alone or in conjunction with others, the most common are multidimensional scaling, clustering, factor analysis, self-organizing maps, and PathfinderNetworks (PFNET).

In our opinion, PFNET with pruning parameters r = ∞, and q = n − 1 is the prime option for eliminating less significant relationships while preserving and highlighting the most essential ones, and capturing the underlying intellectual structure in a economical way.

Although PFNET has been used in the fields of Bibliometrics, Informetrics, and Scientometrics since 1990 (Fowler & Dearhold, 1990), its introduction in citation was due to the hand of Chen (1998, 1999), who introduced a new form of organizing, visualizing, and accessing information. The end effect is the pruning of all paths except those with the single highest (or tied highest) cocitation counts between categories (White, 2001).

The spring embedder type is most widely used in the area of documentation, and specifically in domain visualization. Spring embedders begin by assigning coordinates to the nodes in such a way that the final graph will be pleasing to the eye (Eades, 1984). Two major extensions to the algorithm proposed by Eades (1984) have been developed by Kamada and Kawai (1989) and Fruchterman and Reingold (1991).

While Brandenburg, Himsolt, and Rohrer (1995) did not detect any single predominating algorithm, most of the scientific community goes with the Kamada–Kawai algorithm. The reasons upheld are its behavior in the case of local minima, its capacity to minimize differences with respect to theoretical distances in the entire graph, good computation times, and the fact that it subsumes multidimensional scaling when the technique of Kruskal and Wish (1978) is applied.

We can effortlessly see which are the most important nodes in terms of the number of their connections and, in turn, which points act as intermediaries with other lines, as hubs or forking points.

Whereas factor analysis is a clustering-oriented procedure, PFNET is topology oriented. Yet, they are extremely valuable as complements in the detection of the structure of a scientific domain.

Thus, factor analysis is responsible for identifying, delimiting, and denominating the great thematic areas reflected in the scientogram.

Meanwhile, PFNET is in charge of making the subject areas more visible, grouping their categories into bunches, and showing the paths that connect the different prominent categories, and finally, the overall topology of the domain.

Factor analysis identifies 35 factors in the cocitation matrix of 218 × 218 categories of world science 2002. Through the scree test we extracted 16, which we tagged using the previously explained method; these accumulate 70.2% of the variance (Table 1)

The number of categories included in at least one factor is 195. Twenty-three were not included in any factor (Table 2), and 25 belonged to two factors simultaneously (Table 5).

That is, a category or thematic area occupying a central position in the scientogram will have a more general or universal nature in the domain as a consequence of the number of sources it shares with the rest, contributing more to scientific development than those with a less central position.

The more peripheral the situation of a category or subject area, the more exclusive its nature, and the fewer the sources it will appear to share with other categories; accordingly, the lesser its contribution to the development of knowledge through scientific publications.

An intermediary position favors the interconnection of other categories or thematic areas. 

This broad interpretation of our scientograms not only explains the patterns of cocitation that characterize a domain but also foments an intuitive way for specialists and nonexperts to arrive at a practical explanation of the workings of PFNET (Chen & Carr, 1999).

From a macrostructural point of view, we can distinguish three major zones.

In the center is what we could call Medical and Earth Sciences, consisting of Biomedicine, Psychology, Etiology, Animal Biology & Ecology, Health Care & Service, Orthopedics, Earth & Space Science, and Agriculture & Soil Sciences.

To the right, we can see some other basic and experimental sciences: Materials Sciences & Physics, Applied; Engineering; Computer Science & Telecommunications; Nuclear Physics & Particles & Fields; and Chemistry.

To the left is the neighborhood of the social sciences, with Applied Mathematics, Business, Law, and Economy, and Humanities.

On one hand, it offers domain analysts the possibility of seeing the most essential connections between categories of given domain.

On the other hand, it allows us to see how these categories are grouped in major thematic areas, and how they are interrelated in a logical order of explicit sequences.

2015年4月14日 星期二

Pudovkin, A.I., & Garfield, E. (2002). Algorithmic procedure for finding semantically related journals. Journal of the American Society for Information Science and Technology, 53(13), 1113–1119.

Pudovkin, A.I., & Garfield, E. (2002). Algorithmic procedure for finding semantically related journals. Journal of the American Society for Information Science and Technology, 53(13), 1113–1119.

本研究嘗試利用論文的引用做為參數計算期刊之間的相關因素(relatedness factor),根據計算出來的相關因素找到與目標期刊意義上最相似的期刊。傳統的分類仰賴於根據主觀分析,主觀分析是根據某個或某些特定的分類原因,例如ISI期刊索引報告(Journal Citation Reports,JCR)上的期刊分類便是由經驗法則(heuristic)的主觀方式產生。JCR的作法是在類別建立之後,在同一時間,將新的期刊根據它的相關引用資料進行目測,指定類別;當類別成長,便將類別再細分。除此以外對於個別期刊的分類,也有使用一個未被發表的演算法--Hayne-Coulson algorithm,這個演算法將任何特定的期刊群組做為一個大型期刊(macro-journal),然後產生引用與被引用的期刊資料。在大多數的情況下,這種主觀分析已經足夠,但在一些研究領域中,它被認為是過於粗略而不足並且也受限於與時間的不確定,此外也無法讓使用者可以快速了解哪些期刊是最密切相關的。因此,引進引用索引(citation indexes)與的量化方法被提出來解決這些問題。JCR對每種期刊根據它的引用關係提供了一組最密切相關的期刊,也就是它引用最多的期刊以及引用它最多的期刊,Pudovkin & Garfield (2002)認為這是極為有用並且提供了一種原始的分類,然而由於每種期刊的論文數量不同,使得只能夠得到期刊間關係的淺層感知。因此他們提出了一種期刊間相關因素的測量方式:假定Ri>j表示期刊i和j之間的相關因素,定義Ri>j等於Hi>j * 106 / (Papj * Refi),此處Hi>j是當年度期刊i引用期刊j的次數,Papj與Refi分別是期刊j當年發表的論文數以及期刊i當年論文的參考文獻總數。上述的定義需要注意的是期刊本身的相關因素也許比它對其他期刊的相關因素來得小。此外,為了使兩種期刊A和B之間的相關因素對稱,所以本研究採用RA>B與RB>A中最大的一個,也就是定義RA&Bmax = max(RA>B, RB>A)。本研究以基因與遺傳學領域的核心期刊Genetics為例,研究結果顯示這種根據期刊論文數量加權的相關因素計算方式在發現相關期刊上的效果比未加權的方式來得好,這種方式可以發現原先未被歸入JCR的"Genetics & Heredity"類別但明顯是遺傳學相關的期刊,也可以發現原本歸入這個類別但內容較不相關的期刊。

Using citations, papers and references as parameters a relatedness factor (RF) is computed for a series of journals. Sorting these journals by the RF produces a list of journals most closely related to a specified starting journal.

The method appears to select a set of journals that are semantically most similar to the target journal.

Traditional classification relies on subjective analysis which for one reason or another proves inadequate and is subject to the vagaries of time.

Quantitative methods have been proposed for overcoming these problems. This was greatly facilitated with the introduction of citation indexes in the 1960's and the later introduction of the ISI Journal Citation Reports.

JCR reports inter-journal citation frequencies for thousands of journals. .... Journals are assigned to categories by subjective, heuristic methods.

One of the referees asked for a description of the procedures used by ISI in establishing journal categories for JCR. ... This method is “heuristic” in that the categories have been developed by manual methods started over 40 years ago. Once the categories were established, new journals were assigned one at a time. Each decision was based upon a visual examination of all relevant citation data. As categories grew, subdivisions were established. Among other tools used to make individual journal assignments, the Hayne-Coulson algorithm is used. The algorithm has never been published. It treats any designated group of journals as one macrojournal and produces a combined printout of cited and citing journal data.

In many fields these categories are sufficient but in many areas of research these “classifications” are crude and do not permit the user to quickly learn which journals are most closely related.

JCR provides, for each journal, a set of its most closely related journals based on citation relationships. These are the journals it cites most heavily (cited journals) and also the journals which cite it most often (citing journals). These are extremely useful and provide a crude classification, but unfortunately due to the variations in the sizes of journals one only obtains a superficial perception of the relatedness between two or more specific journals.

We have illustrated the procedure using one core journal in the field of genetics and heredity, the well-known Genetics, published by the Genetics Society of America.

Let journal relatedness of two journals, “i” and “j” be symbolized by Ri>j = Hi>j * 106 / (Papj * Refi), where Hi>j is the number of citations in the current year from journal “i” to journal “j” (to papers published in “j” in all years of ‘j’), Papj and Refi are the number of papers published and references cited in the j-th and i-th journals in the current year.

If we consider a pair of journals, A and B, there may be two indexes: RA>B and RB>A. These can be very different.

It is noteworthy that the citation relatedness of a journal to itself (that is “self-relatedness”) may be lower than its relatedness to some other journals.

Now it is suggested we use the larger of them, RA&Bmax = max(RA>B, RB>A), which we shall call the relatedness factor (RF).

An important feature of the suggested approach is the calculation of SPECIFIC citation relatedness, that is, the new indexes take into consideration the sizes of citing (through the number of references) and cited (through the number of published papers) journals.

The new algorithmic approach enables one to find thematically related journals out of a multitude of journals. ... Weighting citation data by journal size allows identifying journals that are similar in content better than unweighted raw citation data.

In the case of the starting journal Genetics the method identified those journals which are significantly genetic in content, but were not included in the “Genetics & Heredity” category of the JCR. ... Journals included in the “G & H” category are rather heterogeneous in content. Some are highly related to Genetics, while others, as for example journals on medical genetics are poorly related to its content.

JCR has become an established world wide resource but after two or more decades it needs to reexamine its methodology for categorizing journals so as to better serve the needs of the research and library community.

2015年4月9日 星期四

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

本研究比較兩種以內容為基礎的期刊分類以及兩種以演算法為基礎的期刊分類。兩種以內容為基礎的期刊分類分別是ISI的主題分類(Subject Categories)以及Glänzel and Schubert (2003)的領域/次領域分類(field/subfield classification)SOOI,兩種以演算法為基礎的期刊分類則分別是Blondel et al. (2008)提出的展開式(unfolding)社群偵測(community detection)法以及Rosvall, and Bergstrom (2008)的隨機漫步(random walk)矩陣分解(matrix decomposition)法。若是利用以內容為基礎的分類,期刊可以同時指定多個類別;以演算法為基礎的期刊分類則可以使類別內的引用(within-category citation)對類別內的引用(between-category citation)的比率最大化,也就是將期刊彼此之間的引用資料排列成矩陣,經過適當的行列排列後,使得主要對角線(principal diagonal)附近的數值較大,而其他地方則接近0。

各種分類的相關統計資料如表1所示:


由於以內容為基礎的分類方法具有多重分類特性以及以演算法為基礎的分類方法以矩陣分解為目的,從表1上可以觀察到兩種現象:1) 在類別內期刊數的中位數方面,可以看到以內容為基礎的兩種期刊分類方法較以演算法為基礎的期刊分類方法來得多,可配合圖1每個類別期刊數的分佈在0.50上所呈現的情形。另外,圖1也可發現四種分類方法都是對數常態分布(log normal distribution),也就是在這四種分類方法中,相對少數的類別擁有大量的期刊,然而許多類別卻只有少量期刊。並且以演算法為基礎的分類方法比以內容為基礎的分類方法更偏斜(more skewed),也較是上述的情況更嚴重。隨機漫步方法的前十個類別共有57%種期刊,展開方法則有50%,但ISI和SOOI則分別只有15%和31%。


2) 從引用的分布情形來看,兩種以內容為基礎的分類方法的引用次數總計比以演算法為基礎的分類方法多,但隨機漫步方法和展開方法有較多比率分布在類別內,但ISI和SOOI則是主要分布在類別之間。

接下來,以引用式樣(citation patterns)的餘弦相似性(cosine similarity),比較各種分類方法的類別彼此間的相似性。結果ISI和SOOI的中位數分別是0.020和0.066,比隨機漫步方法和展開方法的0.009和0.007高許多,其原因同樣是因為內容為基礎的方法有多重分類的特性,因此類別間的邊緣較模糊,而演算法為基礎的方法在類別間切割得較清楚。然後將各種分類方法的類別依照它們的相似性繪製成網路圖。四種方法繪製的網路圖大致上都可以看出包含兩大群,一個是生物醫學,另一個則是物理學與工程學,兩個大群體透過三個群體相連,包括化學、地理學-環境科學-生態學群體、以及電腦科學,社會科學群體在網路圖上有些分離,透過行為科學/神經科學和生物醫學相連,並且也透過電腦科學與數學和物理學/工程學相連。綜上所述,不同的科學地圖是相似的,但它們在群體內部類別的密度不同。

In this study, we test the results of two recently available algorithms for the decomposition of large matrices against two content-based classifications of journals: the ISI Subject Categories and the field/subfield classification of Glänzel and Schubert (2003).

The content-based schemes allow for the attribution of more than a single category to a journal, whereas the algorithms maximize the ratio of within-category citations over between-category citations in the aggregated category-category citation matrix.

At that time, Leydesdorff & Rafols (2009) were deeply involved in testing the ISI Subject Categories of these same journals in terms of their disciplinary organization. Using the JCR of the Science Citation Index (SCI), we found 14 major components using 172 subject categories, and 6,164 journals in 2006. Given our analytical objectives and the well-known differences in citation behaviour within the social sciences (Bensman,2008), we decided to set aside the study of the (220 − 175 = ) 45 subject categories in the social sciences for a future study.

Our findings using the SCI indicated that the ISI Subject Categories can be used for statistical mapping purposes at the global level despite being imprecise in terms of the detailed attribution of journals to the categories.

In this study, we compare the results of these two algorithms with the full set of 220 Subject Categories of the ISI. In addition to these three decompositions, a fourth classification system of journals was proposed by Glänzel and Schubert (2003) and increasingly used for evaluation purposes by the Steungroep Onderwijs and Onderzoek Indicatoren (SOOI) in Leuven, Belgium. These authors originally proposed 12 fields and 60 subfields for the SCI, and three fields and seven subfields for the Social Science Citation Index and the Arts and Humanities Citation Index. Later, one more subfield entitled “multidisciplinary sciences” was added.

Thus, because research topics are, on the one hand, thinly spread outside the core group and, on the other hand, the core groups are interwoven, one cannot expect that the aggregated journal-journal citation matrix matches one-to-one with substantive definitions of categories or that it can be decomposed in a single and unique way in relation to scientific specialties. The choice of an appropriate journal set can be considered as a local optimization problem (Leydesdorff, 2006).

Citation relations among journals are dense in discipline-specific clusters and are otherwise very sparse, to the extent of being virtually non-existent (Leydesdorff & Cozzens, 2003).

The grand matrix of aggregated journal-journal citations is so heavily structured that the mappings and analyses in terms of citation distributions have been amazingly robust despite differences in methodologies (e.g., Leydesdorff, 1987 and 2007; Tijssen, de Leeuw, & van Raan, 1987; Boyack, Klavans, & Börner, 2005; Moya-Anegón et al., 2007; Klavans & Boyack, 2009).

A decomposable matrix is a square matrix such that a rearrangement of rows and columns leaves a set of square sub-matrices on the principal diagonal and zeros everywhere else.

In the case of a nearly decomposable matrix, some zeros are replaced by relatively small nonzero numbers (Simon & Ando, 1961; Ando & Fisher, 1963). Near-decomposability is a general property of complex and evolving systems (Simon, 1973 and 2002).

The decomposition into nearly decomposable matrices has no analytical solution. However, algorithms can provide heuristic decompositions when there is no single unique correct answer.

Newman (2006a, 2006b) proposed using modularity for the decomposition of nearly decomposable matrices since modularity can be maximized as an objective function.

Blondel et al. (2008) used this function for relocating units iteratively in neighbouring clusters. Each decomposition can then be considered in terms of whether it increases the modularity.

Analogously, Rosvall, and Bergstrom (2008) maximized the probabilistic entropy between clusters by estimating the fraction of time during which every node is visited in a random walk (cf. Theil, 1972; Leydesdorff, 1991).

The data were harvested from the CD-Rom version of the JCR of the SCI and Social Science Citation Index 2006, and then combined. ... The resulting set of 7,611 journals and their citation relations otherwise precisely corresponds to the online version of the JCRs. This large data matrix of 7,611 times 7,611 citing and cited journals was stored conveniently as a Pajek (.net) file and used for further processing.

The 7,611 journals are attributed by the ISI with 11,856 subject classifiers. This is 1.56 (±0.76) classifiers per journal. The ISI staff assign the 220 ISI Subject Categories on the basis of a number of criteria including the journal's title and its citation patterns (McVeigh, personal communication, March 9, 2006; Bensman & Leydesdorff, 2009).

According to the evaluation of Pudovkin and Garfield (2002), in many fields these categories are sufficient, but the authors added that “in many areas of research these ‘classifications’ are crude and do not permit the user to quickly learn which journals are most closely related” (p. 1113).

Leydesdorff and Rafols (2009) found that the ISI Subject Categories can be used for statistical purposes—the factor analysis for example can remove the noise—but not for the detailed evaluation. In the case of interdisciplinary fields, problems of imprecise or potentially erroneous classifications can be expected.

For the purpose of developing a new classification scheme of scientific journals contained in the SCIs, Glänzel and Schubert (2003) used three successive steps for their attribution. The authors iteratively distinguished sets cognitively on the basis of expert judgements, pragmatically to retain multiple assignments within reasonable limits, and scientometrically using unambiguous core journals for the classification. The scheme of 15 fields and 68 subfields is used extensively for research evaluations by the Steunpunt Onderwijs and Onderzoek Indicatoren (SOOI), a research unit at the Catholic University in Leuven, Belgium, headed by Glänzel.

The SOOI categories cover 8,985 journals. Using the full titles of the journals, 7,485 could be matched with the 7,611 journals under study in the JCR data for 2006 (which is 98.3%). These journals are attributed 10,840 classifiers at the subfield level. This is 1.45 (±0.66) categories per journal. One category (“Philosophy and Religion”) is missing because the Arts & Humanities Citation Index is not included in our data. Thus, we pursued the analysis with the 67 SOOI categories.

Using Rosvall and Bergstrom's (2008) algorithm with 2006 data, we obtained findings similar to those of these authors on August 11, 2008. Like the original authors using 6,128 journals in 2004, we found 88 clusters using 7,611 journals in 2006.

Lambiotte, one of the coauthors of Blondel et al. (2008), was so kind as to input the data into the unfolding algorithm and found the following results: 114 communities with a modularity value of 0.527708 and 14 communities with a modularity value of 0.60345. We use the 114 communities for the purposes of this comparison. These categories refer to 7,607 (= 7611 − 4) journals because four of the journals in the file were isolates.

The number of journals per category is log-normally distributed in each of the four classifications. In other words, they all have a relatively small number of categories with a large number of journals and many categories with only a few journals. However, as shown in Figure 1, the classifications based on the random walk and unfolding algorithms are more skewed than the content-based classifications.



Whereas the top-10 categories on the basis of a random walk comprise 57% of the journals (50% for unfolding), they cover only 15% in the ISI decomposition and 31% for the SOOI classification. In the case of skewed distributions, the characteristic number of journals per category can best be expressed by the median: the median is below 30 in the random walk or unfolding classifications, compared with 42 journals for the ISI classification and 141 for the SOOI classification (Table 1).


As presented in the last rows of Table 1, the total numbers of citations in the aggregated matrices based on the ISI or SOOI classifications are much higher because the same citation can be attributed to two or three categories. Thus, whereas random walk and unfolding lead to matrices with most citations within categories (on the diagonal), matrices based on ISI and SOOI classifications lead to matrices with most citations between categories (off-diagonal).

Finally, to measure how similar the categories in the four decompositions are to each other, we computed the cosine similarities in the citation patterns between each pair of citing categories in the four aggregated category-category matrices (Salton & McGill, 1983; Ahlgren, Jarneving, & Rousseau, 2003).

We find again that all the distributions are highly skewed and that the random walk and unfolding algorithms exhibit a much lower median similarity value among categories. The lower medians indicate that the algorithmic decompositions produce a much “cleaner” cut between categories than the content-based classifications.
In conclusion, the analysis of the statistical properties of the different classifications teaches us that the random walk and the unfolding algorithms produce much more skewed distributions in terms of the number of journals per category, but these constructs are more specific than the content-based classification of the ISI and SOOI. The content-based sets are less divided because the boundaries among them are blurred by the multiple assignments.

In summary, although the correspondences among the main categories are sometimes as low as 50% of the journals, most of the mismatched journals appear to fall in areas within the close vicinity of the main categories. In other words, it seems that the various decompositions are roughly consistent but imprecise.

Maps of science for each decomposition were generated from the aggregated category-category citation matrices using the cosine as similarity measure.

The similarity matrices were visualized with Pajek (Batagelj & Mrvar, 1998) using Kamada and Kawai's (1989) algorithm.

The threshold value of similarity for edge visualization is pragmatically set at cosine > 0.01 for the algorithmic decompositions and cosine > 0.2 for the content-based decompositions to enhance the readability of the maps without affecting the representation of the structures in the data.

For the ISI decomposition, the 220 categories (Figure 3) were clustered into 18 macro-categories (Figure 4) obtained from the factor analysis (cf. Leydesdorff and Rafols, 2009).


The map of the SOOI classification was constructed with all is 67 subfields (Figure 5).


Taking advantage of the concentration of journals in a few categories, in the case of random walk and unfolding only the top 30 and 35 categories were used, respectively.


Indeed, the four maps correspond in displaying two main poles: a very large pole in the biomedical sciences and a second pole in the physical sciences and engineering. These two poles are connected via three bridging areas: chemistry, a geosciences-environment-ecology group, and the computer sciences. The social sciences are somewhat detached, linked via the behavioral sciences/neuroscience to the biomedical pole, and via the computer sciences and mathematics to the physics/engineering pole.

As noted above, although categories of different decompositions do not always match with one another, most “misplaced” journals are assigned into closely neighbouring categories. Therefore, the error in terms of categories is not large and is also unsystematic. The noise-to-signal ratio becomes much smaller when aggregated over the relations among categories.

As a second important observation that can be made on the basis of these maps, we wish to point to the differences in category density between the content-based and the algorithm-based maps.

In summary, we were surprised to find that the different science maps are similar except that they differ in the density of categories within groups.

The content-based classifications achieve a more balanced coverage of the disciplines at the expense of distinguishing categories that may be highly similar in terms of journals.

The first finding is that the algorithmic decompositions have very skewed and clean-cut distributions, with large clusters in a few scientific areas, whereas indexers maintain more even and overlapping distributions in the content-based classifications.

Second, the different classifications show a limited degree of agreement in terms of matching categories. In spite of this lack of agreement, however, the science maps obtained are surprisingly similar; this robustness is due to the fact that although categories do not match precisely, their relative positions in the network among the other categories is based on distributions that match sufficiently to produce corresponding maps at the aggregated level.

2015年4月6日 星期一

Chen, C.-M. (2008), Classification of scientific networks using aggregated journal-journal citation relations in the Journal Citation Reports. Journal of the American Society for Information Science and Technology, 59(14), 2296–2304. doi: 10.1002/asi.20935

Chen, C.-M. (2008), Classification of scientific networks using aggregated journal-journal citation relations in the Journal Citation Reports. Journal of the American Society for Information Science and Technology, 59(14), 2296–2304. doi: 10.1002/asi.20935

本研究利用親似傳導法(affinity propagation method, Frey & Dueck, 2007),以彙整的期刊對期刊引用關係(aggregated journal-journal citation relation),對期刊間由相似的引用樣式(citation patterns)形成的科學網路進行分類。過去已有許多以期刊對期刊引用資料進行分析的研究,例如Pudovkin and Garfield (2002) 根據引用資料,發展關係係數(relatedness factor)來發現意義相關的期刊(semantically related journals);Doreian and Fararo (1985)發現網路上結構對等(structure equivalence)的期刊;Leydesdorff and Cozzens (1993)利用主成分分析(principal component analysis)取得科學網路的特徵向量(eigenvectors)。本研究所使用的引用資料包括2001年的SCI(共使用1905種期刊、426065篇文章以及13798138個引用資料)以及2005年的SSCI(共使用1578種期刊、66051篇文章以及2437389個引用資料)。本研究所使用的親似傳導法利用s(i,j)= −dij測量期刊j可以做為期刊i所在類別代表期刊的適合性,而dij的計算為

csij則是期刊間的引用樣式(citation pattern)的相似性:


親似傳導法反覆計算期刊間的兩種數值估算期刊間的代表性,r(i, j)反應期刊j能否代表期刊i的適合程度,

a(i, j)則反應期刊i是否應選擇期刊j作為代表的適合程度,


對期刊i來說,最大的a(i, j) + r(i, j)便指明哪一個期刊j可以代表它。

根據分類的結果,一個分類的專指性(specificity)可以從所有的成員期刊到此分類的代表期刊的平均距離來表示,愈小的平均距離表示這個分類具有愈高的專指性。成員之間的相關性(relatedness of category members)則以所有的期刊之間的平均距離來表示,愈小表示成員間彼此愈靠近。
本研究對SSCI期刊的分類結果共分為23個分類,每一個分類大致符合SSCI的主題分類,然而分類裡所有成員的平均距離比SSCI相對應的分類還要小。

Traditional classification methods (Glänzel & Schubert, 2003) are based on subjective analysis, whose output could vary from one person to another. In other words, these methods are more artistic than scientific.

On the other hand, a quantitative approach to classification is usually constructed based on a set of simple rules, which offers robust classification schemes that do not rely on human interference.

The aggregated journal-journal (J-J) citation data in JCR contain extensive information about interjournal citations, which could provide an understanding of the interaction among various scientific disciplines.

Based on JCR citation data, Pudovkin and Garfield (2002) have used an intuitive criterion (relatedness factor) for finding semantically related journals.

To avoid subjective analysis, various quantitative methods have been proposed to construct a robust classification system of scientific journals using JCR citation information.

A variety of techniques for analyzing J-J citation relationships have been reported in the literature to cluster scientific journals (Doreian & Fararo, 1985; Leydesdorff, 1986; Tijssen, De Leeuw, & Van Raan, 1987).

For example, by applying the notion of structure equivalence to analyze a small set of journals, Doreian and Fararo (1985) have delineated a set of blocks, which contain journals. These blocks have a very close correspondence to a categorization of the journals based on their aims and objectives.

More recently Leydesdorff and Cozzens (1993) have developed an optimization procedure that stabilizes approximated eigenvectors of the scientific network from principal component analysis as representations of clusters. This principal component analysis has been further extended to rotated component analysis (Leydesdorff, 2006; Leydesdorff & Cozzens, 1993), which enables one to focus on specific subsets with internal coherence.

An alternative method of cocitation clustering has been investigated in constructing a World Atlas of Sciences for ISI (Garfield, Malin, & Small, 1975; Leydesdorff, 1987; Small, 1999).

In this article, I propose a quantitative approach to classify the scientific network in terms of aggregated J-J citation relations of JCR using the affinity propagation method (Frey & Dueck, 2007).

The method used by ISI in establishing journal categories for JCR is a heuristic approach, in which the journal categories have been manually developed initially. The assignment of journals was based upon a visual examination of all relevant citation data.

As the number of journals in a category grew, subdivisions of the category were then established subjectively.

Although this is a useful approach, a more robust, convenient, and automatic classification scheme is desired.

The citation data analyzed include the SCI of 2001 and the SSCI of 2005, which are directly computed from the extraction of the CD version of the ISI database.

There are 2,195 journals of impact factor greater than 1 in the 2001 SCI. After removing 290 journals that did not publish any articles in 2001, there are 1,905 journals left in our data set, which contains 426,065 articles and 13,798,138 citations.

For the 2005 SSCI, there are 1,583 journals in the database, of which 1,578 journals have nonzero contents. The SSCI database contains 66,051 articles and 2,437,389 citations.

In principle, the dissimilarity between two journals can be visualized by the differences in their citation patterns. In other words, the citation pattern of each journal is represented by a normalized citation vector, and these vectors form a rescaled citation matrix. The dissimilarity (or similarity) in citation between two journals is related to the scalar product of their citation vectors.

For mapping or visualization, coefficients of similarity are converted into distances such that closely related journals are short distances apart and remotely related journals are long distances apart.

The affinity propagation method takes as input a collection of similarities between journals, where the similarity s(i, j) measures how well journal j is suited to be the representative of a journal category for journal i. Since the goal is to minimize squared error, we set s(i, j) = −dij.

There are two types of messages exchanged between journals, including the responsibility r(i, j), which is sent from journal i to candidate representative journal (RJ) j, and the availability a(i, j), which is sent from candidate representative journal j to journal i. Here the responsibility reflects the accumulated evidence for how well-suited journal j is to serve as the representative for journal i, and the availability shows the accumulated evidence for how appropriate it would be for journal i to choose journal j as its representative.

Taking into account other potential representative journals for journal i, the responsibility is computed iteratively as

where the initial value of a(i, j) is set to zero in the first iteration. Similarly, taking into account the support from other journals that journal j should be a representative, the availability is updated by gathering evidence from journals as to whether each candidate representative would make a good representative journal:

To reflect accumulated evidence that journal j is a representative based on the positive responsibilities sent to candidate representative j from other journals, the self-availability is updated as

During the process of affinity propagation, the sum of availability and responsibility can be used to identify the representative journal of emerging journal categories. In other words, for any journal i, the value of j that maximizes a(i, j) + r(i, j) identifies that journal j is its representative.

In our classifications, the level of specificity of a category can be found by looking at its value of DRJ (the average distance of members of a category to its representative journal), and relatedness of category members is implied by the value of DJ-J (the average J-J distance within a category).

To demonstrate the applicability of the affinity propagation method in clustering a complete data set of journals, we first apply it to cluster journals in the 2005 SSCI database.

Here the cutoff parameter t is set to 0.0001, implying that the maximal value of DJ-J (DJ-Jmax) is 100. This choice of t is quite reasonable since the probability distribution (PD), or normalized histogram (bin size is 1), of DJ-J in the unclustered SSCI journal database is mostly between 0 and 30, as shown in Figure 1.



With a choice of DJ-Jmax = 100, the distance between unrelated journals is much larger than that between related journals. In other words, for any journal category, unrelated journals will not be located in the vicinity of its members (each journal is considered as a point in a high-dimensional space). Thus only correlated journals will be grouped together by the affinity propagation method.

However, if DJ-Jmax is too close to 30, the positions of unrelated journals are not well separated and the distortion to the journal positions due to the introduction of the cutoff would affect the clustering of journals.

For the predicted SSCI classification, only those J-J distances within the same category are considered in calculating its PD of DJ-J.

In Figure 1, there are two peaks observed from the statistical curves of PD in DJ-J, where the first peak shows the relatedness between journals within the database (or categories), while the second peak at DJ-J = 100 indicates the irrelevance between journals within the database (or categories).

For the predicted SSCI classification, clearly its first peak in the PD of DJ-J is much more prominent and the peak width is much more narrow than that of the unclustered SSCI database.

On the other hand, its second peak of irrelevance is much smaller than that of the unclustered database.

The probability distribution of the first peak is found to decrease exponentially with DJ-J, i.e., P = P0 exp[−(DJ-J − d0)/ Δ], where P0 is the peak value, d0 is the peak position, and Δ is the decay width. By fitting the statistical data, we find that d0 = 4 and Δ = 9.08 for the unclustered SSCI curve, while d0 = 2 and Δ = 1.72 for the clustered SSCI curve.

The entire journal set of SSCI is decomposed into 23 journal categories.

The relatedness of journals within a category can be seen as the average value of DJ-J within the category, and the specificity of a category is related to the average distance of category members to its RJ.

For any category, a smaller value of DRJ implies a higher level of specificity, and a smaller value of DJ-J implies that journals within a category are more closely related to each other.

In general most categories in our classification scheme have a corresponding category in the ISI classification scheme, and their value of DJ-J seems to be smaller than that of their counterpart in the ISI classification scheme.

When a larger value of the cutoff parameter is used, the maximal distance of DJ-J becomes smaller. ... Since the high-dimensional J-J distance space is now approximated by a high-dimensional sphere of smaller radius, the resolution in clustering journals is higher in this case. Thus the SCI database is expected to be decomposed into more clusters for t = 10−3, compared to the case of t = 10−4. ... Therefore, from comparing clustering results with different values of the cutoff parameter, the relationship among various disciplines can be revealed.

Our results demonstrate that the affinity propagation method can provide a reasonable classification scheme for either a complete database or an incomplete database. This method does not need the number of categories or their size as an input.

Distance between journals is calculated from the similarity of their annual citation patterns with a cutoff parameter to restrain the maximal distance.

Different values of the cutoff parameter lead to different levels of resolution in the classification of journal network. A more coarse-grained classification is obtained when a smaller value of the cutoff parameter (or a larger maximal J-J distance) is used.

We note that, unlike the ISI classification scheme, which allows overlap in the content of journal categories by subjective decisions, each journal uniquely belongs to a category in our classification scheme.