顯示具有 agglomerative hierarchical clustering 標籤的文章。 顯示所有文章
顯示具有 agglomerative hierarchical clustering 標籤的文章。 顯示所有文章

2016年7月10日 星期日

Janssens, F., Zhang, L., De Moor, B., & Glänzel, W. (2009). Hybrid clustering for validation and improvement of subject-classification schemes. Information Processing & Management, 45(6), 683-702.

Janssens, F., Zhang, L., De Moor, B., & Glänzel, W. (2009). Hybrid clustering for validation and improvement of subject-classification schemes. Information Processing & Management45(6), 683-702.



科學認知映射(cognitive mapping of science)能將科學的結構(the structure of science)加以視覺化,起初應用於資訊服務(information services)、後來發現也可將其應用在科學政策(science policy)與研究評鑑(research evaluation),現在則有越來越多將其應用在發現新興與正在融合中的領域以及主題劃分(subject delineation)的改善。

科學認知映射主要可分為依據引用資訊、依據文本以及混合上述兩種資訊等三種方法。本研究利用文本與引用資訊混合的方法,將2002-2006年Web of Science資料庫內的期刊進行集群,以集群結果產生的認知映射,檢驗目前的期刊主題分類架構,如果可行的話,也將提出改善方式。

本研究首先評估ESI (Essential Science Indicators) 的22個領域主題分類架構,並且將其視覺化。圖1的左右分別是以交互引用與文本方式測量22個ESI領域的Silhouette值,從圖上發現生物學及生物化學(#2)、臨床醫學(#4)、工程學(#7)、植物及動物科學(#19)以及社會科學(#21)等領域上的期刊並沒有足夠好的一致性(coherent)。



從詞語的TF-IDF可以找出每個領域的描述詞語,而這裡也可發現不少領域的描述語互有重疊,例如工程學(#7)與電腦科學(#5)、化學(#3)與材料科學(#11)、植物及動物科學(#19)與環境/生態學(#8),以及生物學及生物化學(#2)、分子生物學及遺傳學(#14)與臨床醫學(#4)等等,此外,從描述社會科學(#21)的詞語也可以了解這個領域高度的異質性(heterogeneity)。


圖2則是以Pajek畫出22個ESI領域的結構圖,圖上也可發現生物學及生物化學(#2)與臨床醫學(#4)、化學(#3)與材料科學(#11)、電腦科學(#5)與工程學(#7)、環境/生態學(#8)與植物及動物科學(#19)等領域之間有很強的關連。

然後將約8300種期刊利用餘弦相似法及Wade的凝聚式階層集群演算法(Wade's  agglomerative hierarchical cluster algorithm)進行集群,再比較集群結果與分類架構。值得說明的是文字部分的資訊可以提供集群結果的標示(labelling),而引用部分則可產生交互引用圖(cross-citation graph)提供視覺化,並且輸入PageRank演算法以決定代表性期刊。

決定集群的數目可以根據集群結果的品質,而集群品質有內在或外在驗證測量(internal or external validation measures)等兩種評估方式。內在驗證只考慮資料與集群的統計特性,例如dendrogram、Silhouette值與模組性(modularity)等;外在驗證需要將集群結果與一個已知的劃分標準進行比較,例如計算兩者間的Jaccard相似性 (Jaccard similarity)。本研究以dendrogram的視覺化方式將期刊首先分為三大群,再分為七群,最後分為22群,三大群約等於自然與應用科學(生物學、農學及環境科學;物理、化學及工程學;數學及電腦科學)、醫學與社會科學以及人文學。以TF-IDF描述語來看,七個群組中有三個屬於自然與應用科學,兩個是生命科學(生物科學及生物醫學與臨床、實驗醫學及神經科學)與兩個是社會科學以及人文學(經濟學、商學及政治學與心理學、社會學及教育學)

圖6是22個群組的結構圖,圖上可看到屬於社會科學以及人文學的群組(#1、#6、#14及#22與#9、#11及#21)、地理學、環境科學、生物學及農學(#2、#15及#19)、物理、化學及工程學(#4、#20及 #5)、數學及電腦科學(#8及#18)、生物科學及生物醫學(#3、#13及#16)與臨床、實驗醫學及神經科學(#7、#10、#12及#17)。



表3比較22個ESI領域與本研究利用引用、文本以及混合等方法產生的22個群組的集群品質,可以發現混合引用及文本資料在各種指標上幾乎都有最好的表現。


圖8是利用Jaccard指標比較集群結果與ESI架構的一致性(concordance)。



最後,本研究並分析期刊轉移(migration)的情形,也就是期刊不屬於原本ESI架構的領域所對應的集群,而被分配到另一個不同集群的現象,好的轉移(Good migration)能使分類的一致性增加,也就是Silhouette值或是模組性增加。本研究希望利用這個現象,從集群與領域的一致性的基礎上提出改善目前期刊主題分類架構的方法。

The main bibliometric techniques are characterised by three major approaches, particularly the analysis of citation links (cross-citations, bibliographic coupling, co-citations), the lexical approach (text mining), and their combination.

A hybrid text/citation-based method is used to cluster journals covered by the Web of Science database in the period 2002–2006. The objective is to use this clustering to validate and, if possible, to improve existing journal-based subject-classification schemes.

In a first step, the 22-field subject-classification scheme of the Essential Science Indicators (ESI) is evaluated and visualised. In a second step, the hybrid clustering method is applied to classify the about 8300 journals meeting the selection criteria concerning continuity, size and impact.

The hybrid method proves superior to its two components when applied separately. The choice of 22 clusters also allows a direct field-to-cluster comparison, and we substantiate that the science areas resulting from cluster analysis form a more coherent structure than the ‘‘intellectual” reference scheme, the ESI subject scheme.

Moreover, the textual component of the hybrid method allows labelling the clusters using cognitive characteristics, while the citation component allows visualising the cross-citation graph and determining representative journals suggested by the PageRank algorithm.

Finally, the analysis of journal ‘migration’ allows the improvement of existing classification schemes on the basis of the concordance between fields and clusters.

The history of cognitive mapping of science is as long as the history of computerised scientometrics itself. While the first visualisations of the structure of science were considered part of information services, i.e., an extension of scientific review literature (Garfield, 1975, 1988), bibliometricians soon recognised the potential value of structural science studies for science policy and research evaluation as well. At present, the identification of emerging and converging fields and the improvement of subject delineation are in the foreground.

The main bibliometric techniques are characterised by three major approaches, particularly the analysis of citation links (cross-citations, bibliographic coupling, co-citations), the lexical approach (text mining), and their combination.

For instance, clustering based on co-citation and bibliographic coupling has to cope with several severe methodological problems. This has been reported, among others by Hicks (1987) in the context of cocitation analysis and by Janssens, Glänzel, and De Moor (2008) with regard to bibliographic coupling. One promising solution is to combine these techniques with other methods such as text mining (e.g., combined co-citation and word analysis: Braam, Moed, & Van Raan, 1991a; combination of coupling and co-word analysis: Small (1998); hybrid coupling-lexical approach: Janssens, Glänzel, & De Moor, 2007; Janssens et al., 2008).

Jarneving (2005) proposed a combination of bibliometric structure–analytical techniques with statistical methods to generate and visualise subject coherent and meaningful clusters. His conclusions drawn from the comparison with ‘intellectual’ classification were rather sceptical.

Despite several limitations, which will be discussed further in the course of the present study, cognitive maps proved useful tools in visualising the structure of science and can be used to adjust existing subject-classification schemes even on the large scale as we will demonstrate in the following.

The main objective of this study is to compare (hybrid) cluster techniques for cognitive mapping with traditional ‘intellectual’ subject-classifications schemes.

In a first study by authors related to the current work, the pilot study of Glenisson, Glänzel, & Persson (2005), further extended and confirmed by Glenisson, Glänzel, Janssens et al. (2005), full-text analysis and traditional bibliometric methods were serially combined to improve the efficiency of the individual methods. It was clear that clusters found through application of text mining provided additional information that could be used to extend and explain structures found by bibliometric methods, and vice versa. However, the integration was still limited to serial combination.

All textual content was indexed with the Jakarta Lucene platform (Hatcher & Gospodnetic, 2004) and encoded in the Vector Space Model using the TF-IDF weighting scheme reviewed by Baeza-Yates & Ribeiro-Neto (1999). Stop words were neglected during indexing and the Porter stemmer was applied to all remaining terms from titles, abstracts, and keyword fields. The resulting term-by-document matrix contained nine and a half million term dimensions (9,473,061), but by ignoring all tokens that occurred in one sole document, only 669,860 term dimensions were retained. Those ignored terms with document frequency equal to one are useless for clustering purposes.

The dimensionality was further reduced from 669,860 term dimensions to 200 factors by Latent Semantic Indexing (LSI) (Berry, Dumais, & O’brien, 1995; Deerwester, Dumais, Furnas, Landauer, & Harshman, 1990), which is based on the Singular Value Decomposition (SVD).

Text-based similarities were calculated as the cosine of the angle between the vector representations of two papers (Salton & Mcgill, 1986).

For simplicity and efficiency, the method used to summarise the subject of a field or cluster is based on selecting the terms with the highest mean TF-IDF weights over all journal papers in the field or cluster, where the IDF factor is calculated on the complete term-by-paper matrix (more than six million papers).

For example, Treeratpituk and Callan (2006) automatically select and assign a few concise labels to hierarchical clusters by combining statistical features from the cluster, parent cluster, and a corpus of general English into a descriptive score.

Geraci, Maggini, Pellegrini, and Sebastiani (2008) label clusters by combining intra-cluster and inter-cluster term extraction based on a variant of the information gain measure, and by looking within the titles of Web pages for the substring that best matches the selected top-scoring words.

The similarities Sij used for clustering were found by calculating the cosine of the angle between the pair of vectors containing all symmetric journal cross-citation values between the two respective journals (i and j) and all other journals (i.e., row or column of the matrix C):

The journal cross-citation graph is also analysed to identify important high-impact journals. We use the PageRank algorithm (Brin & Page, 1998) to determine representative journals in each cluster. Besides, the graph can also be used to evaluate the quality of a clustering outcome.

In order to subdivide the journal set into clusters we used the agglomerative hierarchical cluster algorithm with Ward’s method (Jain & Dubes, 1988).

In general, the number of clusters is determined by comparing the quality of different clustering solutions based on various numbers of clusters. Cluster quality can be assessed by internal or external validation measures. Internal validation solely considers the statistical properties of the data and clusters, whereas external validation compares the clustering result to a known gold standard partition

This compound strategy encompasses observation of a dendrogram, text- and citation-based mean Silhouette curves, and modularity curves. Besides, the Jaccard similarity coefficient is used to compare the obtained results with an intellectual classification scheme.

Up to a multiplicative constant, modularity measures the number of intra-cluster citations minus the expected number in an equivalent network with the same clusters but with citations given at random. Intuitively, in a good clustering there are more citations within (and fewer citations between) clusters than could be expected from random citing.

In Fig. 8, we use the Jaccard index to compare each cluster with every field from the intellectual ESI classification, in order to detect the best-matching fields for each cluster.

Nowadays two ISI systems are widely used, in particular, the ISI Subject Categories, which are available in the JCR and through journal assignment in the Web of Science as well, and the Essential Science Indicators (ESI).

While the first system assigns multiple categories to each journal and is too fine grained (254 categories) for comparison with cluster analysis, the ESI scheme is forming a partition (with practically unique journal assignment) and the 22 fields are large enough. ... This subject-classification scheme is in principle based on unique assignment; only about 0.6% of all journals were assigned to more than one field over a 5-year period.

Fig. 1 presents the evaluation of the 22 ESI fields based on the cross-citation- (left) and text-based (right) Silhouette values (see Section 3.3.3). Since the ESI fields form a partition, this approach allows to evaluate their consistency as if the fields were results of a clustering procedure. Multi-, interand cross-disciplinarity of journals can certainly affect the results.


Several fields seem not to be coherent enough from both perspectives (i.e., the cross-citation and textual approach). Above all, the Silhouette values of field #2 (Biology and Biochemistry), #4 (Clinical Medicine), #7 (Engineering), #19 (Plant and Animal Science) and #21 (Social Sciences) substantiate that at least five of the 22 fields are not sufficiently coherent.

Simultaneously to the above validation, the textual approach also provides the best TF-IDF terms – out of a vocabulary of 669,860 terms – describing the individual fields. These terms are presented in Table 2. Although these terms already provide an acceptable characterisation of the topics covered by the 22 fields, considerable overlaps are apparent between pairs of fields, respectively: Engineering (#7) and Computer Science (#5), Chemistry (#3) and Materials Science (#11), Plant and Animal Science (#19) and Environment/Ecology (#8), as well as Biology and Biochemistry (#2), Molecular Biology and Genetics (#14) and Clinical Medicine (#4). In addition, the terms characterising the social sciences (#21) reflect a pronounced heterogeneity of the field.



The structural map of the 22 ESI fields based on cross-citation links is presented in Fig. 2. For the visualisation we used Pajek (Batagelj & Mrvar, 2003). The network map confirms the strong links we have found based on the best terms between fields #2 and #14, #3 and #11, #5 and #7, and #8 and #19, respectively.

In Table 3 we compare the quality of the partition of 22 ESI fields with the quality of the 22 clusters resulting from citation-based, text-based and hybrid clustering.


The cluster dendrogram shows the structure in a hierarchical order (see Fig. 4). We visually find a first clear cut-off point at three clusters, a second one around seven, and 22 clusters also seemed to be an acceptable/appropriate number.

The number of three clusters results in an almost trivial classification. Intuitively, these three high-level clusters should comprise natural and applied sciences, medical sciences, and social sciences and humanities.

The solution comprising of seven clusters results in a non-trivial classification. The best TF-IDF terms (see Table 5) show that three of these clusters represent the natural/applied sciences, whereas two classes each stand for the life sciences and the social sciences and humanities. This situation is also reflected by the cluster dendrogram in Fig. 4. A closer look at the best TF-IDF terms reveals that the social-sciences cluster (#1 of the 3-cluster solution) is split into the cluster #1 (economics, business and political science) and #6 (psychology, sociology, education), the life-science cluster (#3 in the 3-cluster scheme) is split into clusters #3 (biosciences and biomedical research) and #7 (clinical, experimental medicine and neurosciences) and, finally, the sciences cluster #2 of the 3-cluster scheme is distributed over three clusters in the 7-cluster solution, particularly, the cluster comprising biology, agriculture and environmental sciences (#2), physics, chemistry and engineering (#4) as well as mathematics and computer science (#5).

The social-sciences and humanities clusters form two groups that are each strongly interlinked; one consists of clusters #1, #6, #14 and #22 with focus on humanities, economics, business, political and library science, the other one comprises #9, #11 and #21 with sociology, education and psychology. This is in line with the hierarchical structure shown in Fig. 4. These two groups correspond to the two social-sciences clusters in the 7-cluster solution (cf. Section 4.4).

On the basis of the most important TF-IDF terms (see Table 6) we can assign clusters #2, #15 and #19 to geosciences, environmental science, biology and agriculture, which, in turn, form a larger group corresponding to the first of the three ‘‘megaclusters” in the 7-cluster solution.

These science clusters form two groups, #4, #20 and #5 form one group of chemistry, physics and engineering, while #8 and #18 form the third group comprising mathematics and computer science.

Here we have a biomedical and a clinical group. These two groups are in line with the hierarchical structure of the dendrogram in Fig. 4 but less clearly distinguished in the graphical network presentation (Fig. 6). Nonetheless, the terms provide an excellent description for at least some of the medical clusters: cluster #7 stands for the neuro- and behavioral sciences, #3 for bioscience, #10 for the clinical and social medicine, #13 microbiology and veterinary science, #12 non-internal medicine, #16 hematology and oncology and #17 cardiovascular and respiratory medicine. According to the dendrogram clusters 3, 13, 16 and clusters 7, 10, 12, 17 form one larger cluster each. On the basis of the best terms, we can characterise these groups as the bioscience–biomedical and the clinical and neuroscience group, respectively.

In this subsection we compare the structure resulting from the hybrid clustering with the ESI subject classification. This comparison is based on the centroids of the clusters and fields. The centroid of a cluster or field is defined as the linear combination of all documents in it and is thus a vector in the same vector space. For each cluster and for each field, the centroid was calculated and the MDS of pairwise distances between all centroids is shown in Fig. 7.

In Fig. 8, we use the Jaccard index to determine the concordance between our clustering solution and the ESI Scheme by comparing each cluster with every field, in order to detect the best-matching fields for each cluster. The darker a cell in the matrix, the higher the Jaccard index, and hence the more pronounced the overlap between the corresponding cluster and ESI field.

If clustering algorithms are adjusted or changed, one can observe the following phenomenon. Some units of analysis are leaving clusters they formerly belonged to and end up in different clusters. This phenomenon is called ‘migration’. We can distinguish between ‘good migration’ and ‘bad migration’.

‘Good migration’ is observed if the goodness of the unit’s classi- fication improves, otherwise we speak about ‘bad migration’. We can also apply this notion of migration to the comparison of clustering results with any reference classification. In the following we will use the ESI scheme as reference classification.

Out of 8305 journals under study, there were more than one third, namely, 3204 journals that were not assigned to the cluster which best matches their ESI field. As already mentioned above, we call these journals ‘migrated journals’.

‘Good migrations’ are observed if journals improved their Silhouette values after migration. Based on their titles and scopes (not shown), apparently they should indeed be assigned to the cluster to which they have moved.

Although the Silhouette and modularity values substantiate a more coherent structure of the hybrid clustering as compared with the ESI subject scheme, not all clusters are of high quality. Problems have been found, for instance, in clusters #1 and #12 where interdisciplinarity and strong links with other clusters distort the intra-cluster coherence.


2016年7月7日 星期四

Thijs, B., Zhang, L., & Glänzel, W. (2015). Bibliographic coupling and hierarchical clustering for the validation and improvement of subject-classification schemes. Scientometrics, 105(3), 1453-1467.

Thijs, B., Zhang, L., & Glänzel, W. (2015). Bibliographic coupling and hierarchical clustering for the validation and improvement of subject-classification schemes. Scientometrics105(3), 1453-1467.

Thijs, B., Zhang, L., & Glänzel, W. (2013, January). Bibliographic coupling and hierarchical clustering for the validation and improvement of subject-classification schemes. In Proceedings of ISSI (pp. 237-249).

本研究利用書目耦合(bibliographic coupling)資訊將收錄於Web of Science資料庫內的期刊分群,利用二次相似性(second order similarities)改善書目耦合資訊產生的相似性矩陣過於疏鬆的問題,再以dendrogram及silhouette測量等資訊決定由Ward的凝聚法(Ward’s agglomeration method)產生集群的數目,最後比較集群結果與Glänzel & Schubert (2003)提出的以期刊為基礎的主題分類架構(the journal-based subject-classification scheme),了解兩者間的對應情形,並決定集群的命名。

使用書目耦合的優點是因為需要的資料都已呈現在論文或資料庫上,計算論文(即本文所謂的publications)和期刊間的連結不會有延遲,並且建立連結後將會持續保持一致,不隨時間改變。然而,如同其他使用引用資訊的方法,相關的論文或期刊並無法共同具有所有的參考文獻,在相似性矩陣上無可避免地會有大量的0出現(Janssens, 2007; Janssens et al., 2008),產生極為大量的單一個體(singletons),並影響後續集群分析的品質。過去解決這個問題的一種做法是將引用文獻與詞語相似性的結果混合,例如Janssens et al. (2008);另一種的做法是以二次相似性產生相似性矩陣,例如Janssens (2007)、Ahlgren & Colliander (2009)與 Thijs et al. (2013)等研究。

本研究即是利用二次相似性對Web of Science資料庫內的期刊分群,分析的期刊為2006到2009年間出版100筆論文或以上的期刊,共8282種。作法如下:

1. 以Salton提出的餘弦測量法(cosine measure)產出一次相似(first order similarity),然後以一次相似矩陣再次進行餘弦測量法,產生二次相似性。經過二次相似性的計算後,有10種期刊沒有與其他期刊相連,因此加以移除,因此剩餘的期刊網路上共有8272種。

2. 以Ward的凝聚法產生階層集群,並利用dendrogram及silhouette測量推測可能的集群數目,在本研究由上到下分別有6、14及24種集群數目。silhouette測量方法是對每一種期刊計算一個介於-1和1之間的silhouette數值,正值代表該期刊被分配到適當的集群中。然後將期刊依照其群集分組並以silhouette數值大小排序,產生的圖形可以表示各集群的分群品質,如果在正值的部分有較大的面積,換言之,這個集群具有較多的期刊具有適當的分配,則代表有較好的集群劃分結果。

為了檢驗分群的效果,產生的集群結果與Glänzel & Schubert (2003)提出的以期刊為基礎的主題分類架構進行比較,並且在每一個集群上找出具有代表性的核心期刊(core journals)來分析結果的每個期刊集群。圖五是以網絡來表示集群之間的關係,在圖上可以發現藝術與人文(Arts and Humanities)遠離其他集群,神經科學和行為科學(Neurosciences & Behaviour)介於社會科學和生命科學之間,化學則處於生物科學(Biosciences)、醫學和物理學的中間,較特別的是 一般、區域與社區議題(General,Regional and Community Issues)與生命科學之間有很強的連結。



此外,並且計算14個集群與Glänzel & Schubert (2003)的15個學科(排除Multidiscipline下的期刊)之間的Jaccard指標,呈現為表三。

除了與Glänzel & Schubert (2003)的期刊架構比較以外,本研究也將分群的結果與 ESI (essential science indicators)的類別比較,然而卻發現ESI的劃分與本研究分類結果的結構並不一致。

An attempt is made to apply bibliographic coupling to journal clustering of the complete Web of Science database. Since the sparseness of the underlying similarity matrix proved inappropriate for this exercise, second-order similarities have been used.

Cluster labelling was made on the basis of the about 70 subfields of the Leuven-Budapest subject-classification scheme that also allowed the comparison with the existing two-level journal classification system developed in Leuven. The further comparison with the 22 field classification system of the Essential Science Indicators does, however, reveal larger deviations.

The issue of subject classification and the creation of coherent journal sets has been a major topic in our field since the seventies (see e.g., Narin et al., 1972; Narin, 1976).

The development of computerised methods and the availability of large datasets have shifted the attention from mapping small or single disciplines to the generation of global science maps (Garfield, 1998).

Jarneving (2005) applied bibliographic coupling to map and to analyse the structure of an annual volume of the Science Citation Index.

Janssens et al. (2008; 2009) used a combination of cross-citations and a lexical approach to map journals. Zhang et al. (2010) validated this approach.

The advantage of bibliographic coupling is that there is no delay for the calculation of the link between publications or journals as all data needed are present upon publication or indexing in the database. This also means that link between documents, once established will remain constant over time.

This disadvantage is a result of the very sparse nature of the link matrix (Janssens, 2007; Janssens et al., 2008). The overwhelming number of document pairs does not share any reference at all and thus a large number of zeros occur in the similarity matrix. This deteriorates the quality of the subsequent clustering and may result in an unrealistic large number of singletons (cf. Jarneving, 2005).

As cross-citation data suffers from the same problem, Janssens et al. (2008) introduced a hybrid approach, where they combined citation-based with lexical similarities.

Another solution to overcome the sparseness problem is the use of second order similarities (Janssens, 2007; Ahlgren & Colliander, 2009; Thijs et al., 2013).

A set of journals was compiled from the Web of Science database (SCI-Expanded, SSCI and AHCI). All journals covered in this database between 2006 and 2009 with at least 100 publications in this period are taken into account. This resulted in a set of 8282 journals.

To express the strength of a link between two journals we calculated a first order similarity based on Salton’s cosine measure. The mathematical derivation and interpretation of this similarity measure in the framework of a Boolean vector space model can be found in (Sen & Gan, 1983; Glänzel & Czerwon, 1996).

As bibliographic coupling tends to produce very sparse similarity matrices we applied a second order similarity to reduce this effect. While the first-order similarity is based on the angle between two reference vectors, the second-order similarity is calculated as the cosine of the angle of two vectors holding the first order similarity between two journals.

After the calculation of the second-order similarities, ten journals were removed from the set as they appeared to be singletons without any link to the other journals in the set. The network thus included 8272 journals in total.

Hierarchical clustering with Ward’s agglomeration method was used to create a hard clustering of all the journals.

This method does not provide any automated optimum number of clusters so that the decision was made on the basis of the dendrogram and the silhouette statistics (Rousseeuw, 1987).

Three different levels were chosen. The dendrogram holds strong arguments for a six cluster partitioning while the silhouette plot shows a first peak at 7 clusters. For the highest hierarchical level in the following analysis we use the six cluster solution. At a lower level, the silhouette plot suggests the solutions with 14 and 24 clusters, respectively.

For the evaluation of the specific cluster solution we can rely on the silhouette graphs presented in Figure 4. Each graph presents the silhouette values of the journals in the respective cluster. For each journal a silhouette value is calculated. These values range between 1 and -1 where positive values indicate an appropriate clustering of the journals. Journals are grouped by cluster and ordered from highest silhouette value to lowest. As a consequence the graph gives a good profile of the quality of each cluster. A larger area at the positive side of the vertical axis thus represents a better partitioning.

In order to find an acceptable solution, we decided to use the journal-based subject-classification scheme developed in Leuven (Glänzel & Schubert, 2003). This solution proved most advantageous since both clustering and classification scheme are based on journal assignment. Table 1 presents the hierarchical structure of the three level partitioning. For each cluster the number of journals is mentioned. The labels for the higher levels can be deduced from the lowest level. These labels are taken from the Leuven classification system . The label from the most prominent subject category has been assigned to the corresponding cluster.

Another way to describe the cluster is by using core journals. This notion can be analogously defined as core documents introduced by Glänzel & Czerwon (1996) and extended by Glänzel & Thijs (2011).

In this particular application, a core journal can be identified as journal with at least n links with other journals of at least a given strength r on the second order similarity measure. For the identification of core journals in each cluster we set the number of strong links to at least half the set of journals in the cluster.

As we are using second order similarities this choice is not unreasonable. The value of the strength is chosen such that 12 journals within each cluster comply with both criteria. This means that for more dense clusters the choice of appropriate r-value is higher than in clusters where the journals are not as strongly linked.

Above all, chemistry is at each level a separate cluster. One might expect that at the highest level, chemistry is merged with Physics but we found different patterns.

The second noteworthy observation concerns cluster 17 (Public Health & Nursing). This is a cluster within the ‘Psychology – Neuroscience’ cluster at the highest, six-cluster level. In other partitions or subject classification systems this is attributed to Non-Internal Medicine.

To visualise relations between the 24 clusters we created an additional map. Figure 5 shows these relations.



Despite these multiple assignments we used the Jaccard Index to measure the concordance between the two journal The results are presented in Table 3.

Arts and Humanities is an outlier, Neurosciences & Behaviour acts as a bridge between Social Sciences and Life Sciences, Chemistry takes a central position between Biosciences, Medical Sciences and Physics.Most striking observation in their map is the position of General,Regional and Community Issues which is strongly linked with the Life Science fields.

A 24 cluster solution can be compared with the 22 categories from the classification of Thomson Reuters’ Essential Science Indicators (ESI).

Janssens et al. (2009) showed very low mean silhouette values for the ESI category system in a space with respectively textual distances, cosine similarities of cross-citation vectors and combined distances.

Also in the present study, not all clusters have a unique counterpart in the ESI classification system and vice versa (cf. Janssens et al., 2009). Notably, the ESI fields clinical medicine and engineering, mathematics and social sciences, general are almost uniformly spread over numerous clusters.

Based on this analysis we have to conclude that the segmentation of journals in the ESI categories is not supported by the structure found with bibliographic coupling between journals.

Given this rather weak association between the clustering based on cross-citations and bibliographic coupling, it is a legitimate question to ask which of both methodologies is performing best. A comparison of the mean silhouette values and the silhouette values within each cluster reveals that the methodology presented in this paper results in a more consistent solution.

The 15 cluster solution of cross-citation has a value of 0.04 while the bibliographic coupling results in a value of 0.13.

The main advantage of this method is that clustering can be made as soon as a new database volume is available. The only issue is the lacking cluster labelling that cannot directly be obtained from the method. As a substitute, intellectual classification schemes can be used as reference system. Cluster labelling was made on the basis of the Leuven-Budapest subject-classification scheme that also allowed the comparison with the existing two-level journal classification system developed in Leuven.

The further comparison with the 22 field classification system of the Essential Science Indicators does, however, revealed some striking deviations. These concerned, above all, the fields of clinical medicine, engineering, mathematics and the social sciences. New developments in computer science, neuroscience and psychology as well as in public health (cf. Glänzel & Thijs, 2011) do certainly contribute to such growing deviation.

The main objective of this study was to analyse whether the proposed methodology is appropriate for multi-level journal clustering and to what extent the solutions fit in the framework of traditional subject classification. Further comparison with other solutions such as cross-citation and hybrid methods will be part of future research.

2015年4月21日 星期二

Tseng, Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

Tseng,  Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

近幾十年來,發展出許多科學計量分析技術,包括為了群集(clustering)書目資料所需的各種相似度(similarity)計算技術,如共被引(co-citation)、書目耦合(bibliographic coupling)與詞語共現分析(co-word analysis),這些技術的比較分析可參見Yan and Ding (2012)。並且有很多可以在網路上自由下載使用的軟體工具製作並包裝這些技術,提供科學計量分析應用,知名的軟體工具如CiteSpace (Chen 2006, Chen et al. 2010)、Sci2 Tool (Sci2 Team 2009)、VOSviewer (Van Eck and Waltman 2010)、BibExcel (Persson 2009)及Sitkis (Schildt and Mattsson 2006),這部分的分析則可參見Cobo et al. (2011)。本研究包含兩個部分:提出包含一系列利用書目計量資訊進行群集與映射(mapping)技術的科學計量分析軟體工具集 CATAR,並且將此工具集應用於圖書資訊學(library and information science, LIS)領域後,希望能夠利用期刊群集的結果,確認與分析次領域,以及建議適合研究評估(research evaluation)用途的LIS期刊集合。

Åström (2002)從領域概念的視覺化研究獲得一個結論:期刊的選擇確實影響研究領域如何被知覺與定義,也就是研究領域的界定(delineation)與期刊的選擇有密切關係。已經有許多的研究對圖書資訊學進行次領域界定,而這些研究大多參考ISI的JCR主題分類中與圖書資訊學最相關的類別IS&LS(Information Science and Library Science)。IS&LS類別下並不只包含圖書資訊學的相關期刊,這個類別涵蓋兩個密切相關的領域資訊科學(Information Science)和圖書館學(Library Science),此一範圍與圖書資訊學有些微不同。根據Leydesdorff (2008),JCR主題分類以期刊的題名、引用模式(citation patterns)等等做為標準進行分類,但是這個分類結果與從資料庫本身的引用資料所產生的網路上的主要成分(principal components)得到的分類結果並不十分相符。因此次領域界定研究大多經過人為的挑選做為分析資料的期刊,並沒有完整收錄IS&LS主題下的所有期刊。

進行次領域界定時常使用的技術包括:利用共被引分析比較一對項目,利用凝聚式階層群集(agglomerative hierarchical clustering, AHC)將項目分群產生樹狀圖(dendrogram),利用多維尺度(multi-dimensional scaling, MDS)產生視覺化的二維或三維映射圖。若干重要的研究如:Åström (2002)從圖書資訊學重要期刊中選取1135篇出版在1998到2000年的文章,利用BibExcel軟體工具進行作者共被引(author co-citation)以及關鍵詞共現分析,並產生MDS映射圖,52位高被引作者的共被引產生三個群集:"硬"資訊檢索(hard information retrieval)、"軟"資訊檢索(soft information retrieval)以及書目計量學(bibliometrics),47個較常出現的關鍵詞則分為圖書館學(library science,LS)、資訊檢索(information retrieval,IR)及書目計量學。Åström (2002)認為作者共被引分析沒有出現圖書館學的原因可能與圖書館學研究的出版管道有關,如果引用的資料像是書籍或地區期刊沒有出現在JCR,圖書館學作者便無法出現在引用為基礎的排名上。Åström (2007)對55種在JCR 2003主題類別下的期刊,選擇21種圖書資訊學相關期刊的13605篇文章進行文件共被引分析,在從1990到2004年的三個時段發現圖書資訊學可分為資訊計量學(informetrics)和資訊搜尋與檢索(information seeking and retrieval)兩個穩定的次領域,而隨著全球資訊網的普及,網路計量學(webometrics)在兩個次領域上都成為主要的研究議題。Jassen et al. (2006) 對2002到2004年五種圖書資訊學相關期刊的938篇文章,應用一系列的全文分析技術以及MDS和AHC,將938篇文章分為六個群集:兩個群集與書目計量學有關、一個群集為IR、一個包含一般議題、另兩個較小但愈來愈重要的群集分別是網路計量學和專利分析(patent analysis)。Moya-Anegon et al. (2006)從24種較有影響力的期刊中選擇17種期刊,排除將資訊科學(information science, IS)應用到特定技術或知識領域(例如:醫學、地理學、電訊傳播等),從17種期刊引用的參考文獻,對77位最常被引用的作者和73篇最常被引用的期刊進行共被引分析,映射使用的技術包括MDS和AHC以及自組織映射圖(self-organizing map)。作者共被引分析的結果產生六個次領域:科學計量學、引用分析、書目計量學、"軟"(認知導向)資訊檢索、"硬"(演算法導向)資訊檢索以及傳播理論(communication theory)。而期刊共被引分析的結果則有四個群集:IS、LS、科學研究(science studies)以及管理學(management)。在期刊共被引分析的科學研究大致上可以對應為作者共被引分析的科學計量學、引用分析、書目計量學,IS為"軟"資訊檢索和"硬"資訊檢索。如Åström (2002)同樣的原因,LS沒在作者共被引分析的結果當中。Waltman et al. (2011)以JASIST為種子,選擇與該期刊共被引較多的期刊,連JASIST共48種,進行期刊的書目耦合(bibliographic coupling)分析,並且利用VOSviewer呈現視覺化結果,共分為LS、IS以及科學計量學等3個次領域。Milojevic et al. (2011)使用詞語共現分析探討1998到2007年出版的16種期刊上的10344篇文章,16種期刊根據Nisonger and Davis (2005) 的研究所挑選,分析100個文章題名上最常出現的詞語,進行共現分析,並以AHC歸類,結果三個主要群集為LS、IS以及書目計量學/科學計量學。

Åström (2002)以關鍵詞的共現分析所得到的結果包括LS次領域,但作者共被引分析所得到的映射圖上並沒有產生這個次領域。Moya-Anegon et al. (2006)的期刊共被引分析與作者共被引分析也略有不同,期刊共被引分析的結果上有作者共被引分析沒有的LS和管理學兩個次領域,反之,作者共被引分析的結果上則可以發現期刊共被引分析沒有的傳播學理論(communication theory)。一般認為這和作者引用的行為有關,LS作者的引用次數大多沒有達到分析的門檻,因此無法在上述兩個研究的作者共被引分析結果上呈現。

Ni et al. (2012)從JCR的IS&LS類別下的61種期刊,排除3種非英語的期刊,將選取的58種期刊進行場域-作者耦合(venue-author coupling)、期刊共被引分析、詞語共現分析、期刊連結(journal interlocking)等四種分析。分析的結果再進行MDS與AHC分析,四種方式所得到一致的次領域包括:管理資訊系統(managment information systems, MIS)、IS、LS和特殊化群集(specialized clusters),並且在四種方法所得到MDS映射的圖形上都可以發現MIS與其他群集分離,Ni and Ding (2010)與Ni and Sugimoto (2011)建議JCR上的圖書資訊相關期刊應進行適當的重組。

本研究(Tseng and Tsai 2013)應用的資料範圍為2000到2004與2005到2009在Web of Science 的Journal Citation Report中 Information Science & Library Science (IS&LS)主題分類下的所有期刊,在前期(2000~2004年)共50種,後期(2005~2009年)共66種。本研究的分析程序採用Borner et al. (2003)整理的一般工作流程,步驟包括:1) 資料蒐集(data collection);2)文本分段(text segmentation);3)相似性計算(similarity computation);4)多階段群集(multi-stage clustering);5)群集標名(clustering labeling);6)視覺化(visualization);7)面向分析(facet analysis)。這些步驟中所需的技術都已經整合到軟體工具CATAR(Content Analysis Toolkit for Academic Research, http://web.ntnu.edu.tw/~samtseng/CATAR/)上。在計算文件間的相關性時,本研究以一種期刊做為一個文件,所有論文引用的期刊做為文件的特徵,然後利用Dice係數(Salton 1989)計算期刊相似性,例如兩種期刊X與Y,R(X)與R(Y)分別是它們引用的期刊,它們之間的相似性計算為Sim(X, Y) = 2 ∙ |R(X)∩R(Y)|/(|R(X)|+R(Y)|)。也就是利用書目耦合計算期刊之間的相似性。期刊的群集則是利用完全連接階層群集法(complete-linkage hierarchical clustering)。首先將每個文件視為一個群集,然後將一對最相似的群集合併起來,產生一個較大的群集,然後重複進行上面的步驟,而兩個群集的相似性定義為兩個群集間最小的文件相似性,如果相似性超過某個預先設定的閾值,便將兩個群集合併,一直到無法再產生合併為止。此外,本研究採用Silhouette指標(Ahlgren and Jarneving 2008; Rousseuw 1987; Jassen et al. 2006)。

此一研究的資料包含JCR的IS&LS主題下的期刊,分為2000-2004年與2005-2009年兩個時期,前一個時期包含50種期刊,9546筆論文資料;後一時期則有66種期刊,11471筆論文資料。從群集結果的樹狀圖(dendrogram)和MDS映射的結果顯示,IS&LS主題下的期刊在兩個時期都有IR、MIS、科學計量學、學術圖書館(academic library)、醫學圖書館(medical library)、館藏發展(collection development),以及開放取用(open access)和地區圖書館(regional library)兩個後期出現並且較小的群集。並且MIS群集的期刊在知識基礎(intellectual base)上與IS&LS主題的其他期刊分離,表示這群集下的期刊具有較特殊的引用模式。本研究以期刊的書目耦合進行分析,從期刊知識基礎(intellectual base)得到MIS群集與其他分離的研究結果,與Ni et al. (2012)利用期刊共被引分析、期刊連結、術語使用(terminology usage)和合著(co-authorship)研究等不同方法的研究結果相同,這也為許多探討圖書資訊學認知結構的研究認為不應將MIS相關期刊與其他期刊包含在ISI的同一個主題IS&LS下,在進行分析時需要排除MIS相關期刊提供了佐證(Larivière et al. 2012)。此外,並且以多樣性指標(diversity index)分析群集特性,揭露出某些次領域具有地區(regional)特性。

2015年4月2日 星期四

Wang, F., & Wolfram, D. (2014). Assessment of journal similarity based on citing discipline analysis. Journal of the Association for Information Science and Technology.

Wang, F., & Wolfram, D. (2014). Assessment of journal similarity based on citing discipline analysis. Journal of the Association for Information Science and Technology.

利用Web of Science的主題分類,計算引用期刊的學科頻率分布能夠提供被引用期刊進行相似性比較的特徵,相較於共被引方法,這種相似性比較的維度較小,可以減少許多計算量。本研究比較Web of Science的資訊科學與圖書館學主題分類下的40種高影響力期刊,並以多維尺度法(multidimensional scaling)和階層式群集分析(hierarchical cluster analysis)比較比較所提出的方法與共被引方法的相似性估算結果。分析期刊的出版時間範圍為1987到2011,以5年為一個時期進行分析。在各期刊中,以Scientometrics (SCI)以及Journal of the Association for Information Science and Technology (JASIST)的引用期刊分布的學科最多元,因為JASIST有較廣的涵蓋範圍以及其他領域都對測量研究(metrics research)感到興趣。產生的映射圖與群集結果顯示某些期刊並不接近其他期刊。相似性估算結果顯示引用學科分析與共被引分析相似,各個時期兩種方法所得到的結果在分為三個群集的情況下,大多可以發現包含一個LIS群集、一個MIS群集以及一個較分散而邊緣的群集,不過組成群集的成員也有些不同,因此Wang and Wolfram (2014)建議可以引用學科分析做為共被引分析的補充。

The frequency distribution of disciplines by citing articles provides a signature for a cited journal that
permits it to be compared with other journals using similarity comparison techniques.

As an initial exploration, citing discipline data for 40 high-impact-factor journals assigned to the “information science and library science” category of the Web of Science were compared across 5 time periods. Similarity relationships were determined using multidimensional scaling and hierarchical cluster analysis to compare the outcomes produced by the proposed citing discipline and established cocitation methods.

The maps and clustering outcomes reveal that a number of journals in allied areas of the information science and library science category may not be very closely related to each other or may not be appropriately situated in the category studied.

The citing discipline similarity data resulted in similar outcomes with the cocitation data but with some notable differences. Because the citing discipline method relies on a citing perspective different from cocitations, it may provide a complementary way to compare journal similarity that is less labor intensive than cocitation analysis.

The application of visualization techniques to groups of bibliographic entities (publications, journals, or authors) provides a method for assessing the closeness of relationships among entities of interest. ... On a fundamental level, these investigations allow us to understand better the structure of disciplines based on the production of scholarship and how this changes over time (e.g., White & McCain, 1998). On a more specific level, findings can help to assess the impact of entities of interest or to situate disciplines or specializations within a larger context.

Leydesdorff and Cozzens (1993) studied how to delineate and attribute journals to specialties based on journal−journal citations and their changes over time. They demonstrated how the data could be used to construct macrojournals, consisting of aggregations of journals around a central journal.

Pudovkin and Garfield (2002) developed a journal relatedness factor based on citing and cited journals. The method was proposed to help identify thematically related journals.

Similarly, Glänzel and Schubert (2003) proposed the categorization of journals using a three-step process involving predefined categories, journal classification, and article classification for articles in journals with ambiguous subject assignments based on references.

Rafols and Leydesdorff (2009) compared the outcomes of two algorithms for the decomposition of large matrices against Web of Science (WoS) subject categories and Glänzel and Schubert’s categorization. The four methods resulted in similar map outcomes on a large scale.

Leydesdorff and Schank (2008) visualized and animated the disciplinary ties of three seed journals over time to demonstrate relationships among journals and their interdisciplinarity.

Boyack and Klavans (2010) compared results from cocitation analysis, bibliographic coupling, direct citation, and a hybrid approach for accuracy of outcomes in representing research fronts for a large corpus of biomedical literature. They noted that bibliographic coupling performed the best in representing the research fronts.

White (2000) proposed the use of citers to identify characteristics of a given author’s research such as an author’s citation identity, which consists of all the authors a given author cites. White also introduced the idea of citation image-makers, consisting of the authors who refer to a cited author. The citation image-makers approach may also be applied to journals, where citing authors constitute the citation image-makers of the journal.

Yan, Ding, Milojević, and Sugimoto (2012) explored community structures in IR research by combining topic modeling and community detection with IR literature to reveal the changing landscape of IR research.

To reduce the dimensionality of the similarity comparison, disciplinary identifiers for citing articles/journals may be used to reduce the number of comparisons that have to be made.

For the purposes of this study, WoS research areas are used. In this paper the research areas are referred to as disciplinary assignments.

This research is guided by several questions.
1. Does the frequency distribution of disciplines of citing journals permit comparison of journal similarities in a meaningful way?
2. Are the results of such a comparison similar or complementary to the better-established approach of cocitation analysis?
3. Do the similarities among journals within the same disciplinary categorization change over time as reflected in the changes in the frequency distribution of citing journal disciplines?
4. Can these similarities (or distances) provide decision support for whether journals should be grouped together in citation index services such as Thomson Reuters’ Journal Citation Reports?

Forty high-impact journals included in the Thomson Reuters’ 2011 Journal Citation Reports grouped in the category ISLS were selected for the study.  ... In addition to many of the journals rated highly in library and information science (LIS), as evidenced by a perception study of LIS deans and Association of Research Library directors conducted by Nisonger and Davis (2005), this category includes journals in allied areas such as management information systems (MIS), geographic information systems, and medical informatics.

Among the 20 highest-impact journals listed in the ISLS category, only 3 are included in the top 20 journals rated by LIS deans based on their familiarity with these journals. The majority of the remaining journals in the top 20 based on impact factor could be argued to be from allied areas given their additional classification in other WoS research areas and the lack of familiarity or resulting lower prestige as determined by LIS deans.

Citing article/journal data were collected from 1987 to 2011 and were divided into 5-year intervals.

For each journal, all articles, review articles, and conference proceeding articles were selected; all other publication types such as cited material were excluded. For each time period, the “create citation report” in the WoS was selected to identify all citing articles. The number associated with “citing articles” was then selected to retrieve the list of citing articles. The WoS “analyze results” feature was next selected for the list of citing articles. On the results analysis page, “research areas” were selected as the ranking field to provide the tabulated list of citing disciplines. The ranked list of citing disciplines was then copied into an MS Excel spreadsheet.

The list of research areas and their frequencies represent the journal’s citing discipline profile for each time period.

Salton’s cosine similarity measures were calculated for each pair of journals to produce a symmetric
matrix of journal similarity values ranging between 0 and 1 (Ahlgren, Jarneving, & Rousseau, 2003, 2004; Egghe & Leydesdorff, 2009; Leydesdorff, 2006, 2007) for each time period.

To provide a baseline comparison, a cocitation analysis was also conducted with the same journals.

Multidimensional scaling (MDS) PROXSCAL analysis and hierarchical cluster analysis in SPSS v.20 were applied to the symmetric similarity matrices.

The PROXSCAL algorithm was used instead of ALSCAL for the MDS procedure because it allows similarity or dissimilarity matrices to be used and has been shown to provide superior results for cocitation studies (Leydesdorff & Vaughan, 2006).

For hierarchical clustering, Ward’s method was used. Minkowski distance and squared Euclidean distance were each explored and produced the same outcomes at the three-cluster level. Clustering outcomes were superimposed onto the MDS maps.

Library Resources and Technical Services (LRTS) consistently attracted citations from the fewest discipline areas, indicating a narrower interdisciplinary focus. In fact, the number of citing article disciplines has declined over the past decade for this journal, possibly indicating even narrower interdisciplinary impact.

Scientometrics (SCI) and the Journal of the Association for Information Science and Technology (JASIST), on the other hand, at different time periods each attract the most disciplinarily diverse citations. These outcomes are not unexpected given the broad coverage of JASIST and the interest in metrics research by other disciplines.



The MDS map of the journals using the proposed citing discipline approach for the first period appears in Figure 1. Among the journals, 14 of the 22 are situated in close proximity. A secondary group with three journals is situated on the periphery.

In combination with the cluster-analysis groupings, one can see at the three-cluster level that the tightly clustered journals are core to LIS.

A more widely dispersed second cluster of five journals consists of LIS and allied area journals in MIS. ... It is interesting to note that Government Information Quarterly (GIQ), International Journal of Geographical Information Science (IJGIS), and Journal of the Medical Library Association (JMLA)—at the time, still the Bulletin of the Medical Library Association–are situated more closely to and are clustered with the journals associated with the MIS area.

A peripheral “Other” cluster contains three journals.  ... Telecommunication Policy (TP), Journal of Scholarly Communication (JSP), and Social Science Information (SSI) are situated on the periphery of the map for the first and second time periods, indicating little similarity with the other journals in the citing discipline distributions.


The equivalent cocitation analysis map (Figure 2) at the three-cluster level, produces similar outcomes, but with several notable differences.

The International Journal of Information Management (IJIM) is situated more closely to LIS journals than to those in MIS.

Two of the MIS journals are situated in their own cluster along with GIQ and TP, equivalent to the “other” category. IJGIS appears at the periphery of the map in the MIS category.

The remaining journals are subdivided into two clusters that may be characterized broadly as information science and library science, respectively, with JSP and SSI being a part of these clusters.

There is a 63.6% overlap (14 of 22 journals) in the cluster assignments, indicating that there is a moderate level of agreement between the two approaches.

For the second time period, the three clusters for the citing discipline-based analysis consisted of a group of 12 journals representing the LIS area, an emerging cluster of journals focusing on the MIS area and several journals in allied areas, and an “other” group consisting of JSP, SSI, TP, and IJGIS.

The cocitation analysis outcomes for the second time period reveal a similar mapping arrangement and clustering of journals, with 15 journals corresponding to the LIS category, eight representing a group with an MIS focus, and an “other” category consisting of journals in allied areas.

The citing discipline MDS map and cluster analysis results for the three-cluster level are similar to the first two time periods, but with more distinctive LIS, MIS, and other clusters as the number of journals in each cluster has grown.

The cocitation analysis map and resulting clusters at the three-cluster level consist of primarily LIS journals, those in MIS, and the other category similar in composition to the citing discipline outcome. ... The cluster assignment match at the three-cluster level between the citing discipline and cocitation analysis methods is 88% (29 of 33 journals), indicating a high level of agreement.

The citing discipline MDS map for the fourth time period is similar to that for the previous time period.

Of note with the cocitation cluster analysis outcome for the fourth time period is a much larger other category that includes a number of journals categorized as LIS by the citing discipline method. GIQ and INFSOC are situated between the LIS and MIS groups, although they are placed in the other group.

The citing discipline and cocitation maps for the fifth time period appear in Figures 5 and 6, respectively. The outcomes for the citing discipline approach are quite similar to those for the third and fourth time periods, with well-defined LIS and MIS categories and a more scattered other category on the periphery.

The three clusters based on the cocitation analysis data again reflect the LIS, MIS, and other groupings. There are fewer members in the other category than for the fourth time period

Much in the same way that dimensionality reduction used in certain statistical methods and IR allows for simplified comparisons, the use of the WoS research areas by citing journals and their frequency instead of citing authors or citing journals provides a less computationally intensive way to assess journal similarity by reducing the dimensionality of the comparisons and the computational overhead.


2014年8月15日 星期五

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words. Journal of the American Society for Information Science and Technology, 62(10), 1933-1953.

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words.Journal of the American Society for Information Science and Technology,62(10), 1933-1953.

Scientometrics

圖書資訊學(LIS)為對於記錄下來的資訊(recorded information)和具有文化意義的文物與標本(culturally meaningful artifacts and specimens)有興趣的研究領域(Bates, 2010),包括的領域有檔案學(archival science)、 書目(bibliography)、文獻與文類理論(document and genre theory)、資訊學(informatics)、資訊系統(information systems)、知識管理(knowledge management)、圖書資訊學(LIS)、博物館研究(museum studies)、記錄管理(records management)和資訊的社會研究(social studies of information)。過去有許多研究嘗試定義與描述圖書資訊學的領域並且確認其中包含的研究主題,這些研究使用的方法相當廣泛,包含Järvelin & Vakkari (1990, 1993)採用內容分析(content analysis);Åström (2007, 2010)、Moya-Anegón, Herrero-Solana, & Jiménez-Contreras (2006)和 Persson (1994) 針對期刊或期刊文章進行書目計量分析 (bibliometric analysis) ; Moya-Anegón et al., (2006)和White & McCain (1998)針對作者進行書目計量分析 ;Åström (2002)、 Ding, Chowdhury, & Foo (2001) 和 Janssens, Leta, Glänzel, & De Moor (2006)利用從題名、摘要或全文抽取的詞語進行詞語的共現分析(co-word analysis) ;Sugimoto & McCain (2010)則是用索引詞語的三元共現分析(tri-occurrence analysis) ; van den Besselaar & Heimeriks (2006)利用詞語和參考文獻的組合進行分析;以及Sugimoto, Li, Russell, Finlay, & Ding, (2011)和 Sugimoto & McCain (2010)所使用的主題模型分析方法。

上述的這些方法,許多必須依賴於作者對於領域知識的了解,才能了解領域的主題與認知結構(cognitive structure),例如White & McCain (1998)基於最重要的作家的集群,觀察資訊科學由圍繞在一個微弱中心的許多專業所組成;Åström (2010)則是透過作者與期刊的映射圖說明這個領域的圖書館學(LS)和資訊科學(IS)之間具有差距。除了是認知結構較不直接的指標之外,引用分析另一個的問題是不同的次領域有不同的發表與引用實務。

論文題名包含許多能夠指出該文章內容的詞語(Buxton & Meadows, 1977; Meadows, 1998)。因此,本研究採用的方法是利用期刊論文題名上的重要詞語進行分析。分析的資料來自16種LIS期刊於1988到2007年發表的10344筆論文資料。

選取100個最常出現於題名的詞語。

本研究使用的分析技術包含詞語的相對頻率(relative frequency)並且根據詞語的共現進行叢集,最後並將詞語以及期刊與發表年度等進行多維尺度分析(multidimensional scaling, MDS),產生視覺化的結果。

詞語的共現分析以及階層式集群分析的結果發現三個主要分類LS(圖書館學)、IS(資訊科學)、SCI-BIB(科學計量學-書目計量學)以及兩個較小的分類資訊尋求行為(information-seeking behavior)和書目指導(bibliographic instruction)。LS可再細分為學術圖書館專業(academic librarianship)、公共圖書館專業(public librarianship) (包含館藏建立)、資訊素養和學校圖書館專業(information literacy and school librarianship, technology)、政策(policy)、全球資訊網(the web)、知識管理(knowledge management)、數位圖書館(digital libraries)、電子商務(e-commerce)、法律(law)以及學術出版(scholarly publishing)等主題。IS則包含資訊檢索(information retrieval)、網路搜尋(web search)、分類目錄(catalogs)以及資料庫(database)等主題。SCI-BIB也有書目計量指標(bibliometric indicators)、作者生產力(author productivity)與引用研究(citation study)等主題。整體的結構如下圖

從詞語的使用可以發現LIS中有某些持續出現的核心詞語,但也有一些詞語的使用在20年間有明顯的變化,這些都是與科技相關的(technologically related)詞語,這個現象符合Saracevic(1999)所宣稱的LIS是個科技驅動的(technology driven)領域。大致上來說,LIS內的改變可以從資料庫(database),到數位圖書館(digital libraries),到全球資訊網(the World Wide Web)等詞語使用的移轉上看得出來。

除了科技驅動的特徵外,LIS同時也有很大的範圍在討論資訊尋求行為,這是LS和IS都共同關心的課題。

A number of empirical studies of LIS have been conducted with the aim of describing and defining the field and identifying research areas within it. These studies applied a wide array of approaches: content analysis (Järvelin & Vakkari, 1990, 1993); bibliometric analysis of journals and journal articles (Åström, 2007, 2010; Moya-Anegón, Herrero-Solana, & Jiménez-Contreras, 2006; Persson, 1994); bibliometric analysis of authors (Moya-Anegón et al., 2006,White & McCain, 1998); co-word analysis of both index terms and words extracted from titles, abstracts, and full text (Åström, 2002; Ding, Chowdhury, & Foo, 2001; Janssens, Leta, Glänzel, & De Moor, 2006); tri-occurrence analysis of index terms (Sugimoto & McCain, 2010); analysis of word-reference combinations (van den Besselaar & Heimeriks, 2006); and topic analysis (Sugimoto, Li, Russell, Finlay, & Ding, 2011; Sugimoto & McCain, 2010).

Some notable studies of cognitive structure of LIS have interpreted topics post hoc, by assigning topicality based on knowledge of the author’s domain (e.g., White & McCain, 1998). In White and McCain’s influential visualization of LIS, they concluded that “information science lacks a strong central author, or group of authors, whose work orients the work of others across the board. The field consists of several specialties around a weak center” (p. 343). However, this analysis was based foremost on the clustering of authors, rather than topics. Similarly, Åström (2010) examined the divide between LS and IS components of the field by a bibliometric mapping of authors and journals. Topicality was assigned through expert knowledge of the domains in which these authors wrote and journals published.

Of the various components of textual documents, the titles, and the choice of words in them, are of particular importance. Title words function as “attention triggers” (Bazerman, 1985, 1988). They are devices for capturing interest in the world where information overload is a norm. Title words
have been called “signal-words”1 (Rip & Courtial, 1984) and “macro-actors” or “macro-terms”2 (Callon et al., 1983). Titles of journal articles themselves have undergone a change during the 20th century, becoming more informative, more specific, and containing a larger number of words that indicate article content (Buxton & Meadows, 1977; Meadows, 1998). Leydesdorff (1989) claims that “title words seem to offer a means of making visible the internal cognitive structure” (p. 217) of a discipline. He also claims that “word structure reflects internal intellectual organization in terms
of the codification of word usage in the relevant disciplines” (Leydesdorff, 1989, p. 221). 

Co-word analysis is based on co-occurrence of words (all words, or selected keywords) extracted from titles, abstracts, or text in general, or the index terms assigned by authors or indexers. Co-word analysis is a method that derives “higher level structures from word-occurrence patterns in text” (Chen, 2003, p. 139). Of particular importance in the context of this study is that co-word analysis is “a means to the elucidation of structures of ideas, problems, and so on, represented in appropriate sets of documents” (Whittaker, Courtial, & Law, 1989, p. 473). 

Although co-word analysis has its limitations, (e.g., Leydesdorff, 1997) primarily because of the
change of usage and meaning of words and the lack of context, such analysis has been considered particularly useful in tracking the development of scientific fields over time (Callon et al., 1991; Noyons & van Raan; Rip & Courtial, 1984), which represents another goal of this study.

Although citation analysis is not subject to the same limitation, it is a less direct indicator of cognitive structure. As already mentioned, studies using citations require post hoc assignment of topics. In addition, citation analysis of LIS is less effective in analyzing the cognitive structure of entire fields due to the different publication and citation practices of subfields, thus leaving even large subfields such as LS often invisible.

Selection of journals and articles. Articles from 16 LIS journals were chosen for inclusion in this study. The journals were selected from a ranked list of the most important journals in the field, according to deans and directors of American Library Association (ALA)-accredited, MLS programs in North America (Nisonger & Davis, 2005).

From this journal set, all research and review articles (10,344) published between 1988 and 2007 were included in the analysis.

Identification of the most frequently occurring LIS words and phrases. Word frequency is an important measure in content analysis. This measure is used to identify the most important research topics or concepts in a field by focusing on the most frequently occurring words.

In this study, we base all analyses on the 100 most frequently occurring LIS words or phrases. 

2013年12月2日 星期一

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

科學映射圖(scientific mapping)能夠科學結構(scientific structure)視覺化,幫助使用者確認科學主題(scientific themes)並從而發現新知識的有用工具之一。過去的研究曾經使用過作者、文章與等映射單位。在計算映射單位之間的關連,Börner, Chen, and Boyack (2005) 將關連性的測量方法(relatedness measures)分為引用連結(citation linkages)與共現相似性(co-occurrence similarities)等兩大類,而本研究則將目前常用來評估作者間的關連分為直接引用(direct citation)、共被引分析(cocitation analysis)、合著分析(co-authorship analysis)、書目耦合分析(bibliographic coupling analysis)以及共詞分析(co-word analysis)等五種方法。也有研究以發展出整合文字內容與連結的測量方法來計算期刊(Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005)與文章(Liu et al., 2010)間的關連。本研究建議兩種以詞語為基礎並利用向量空間模式(vector space modeling)的方法和另一種基於LDA(latent Dirichlet allocation)的主題模型方法來測量作者之間的關連。本研究將第一種方法稱為靜態(static)的特徵,以每位作者曾寫過的論文內容為基礎產生代表這位作者的特徵向量,也就是代表這位作者的特徵向量是所有他寫過的論文的特徵向量總和,任何兩位作者之間的關連是對應於他們的作者特徵向量之間夾角的餘弦值(cosine value)。第二種方法則是動態(dynamic)的特徵,如果兩位作者之間沒有合著的論文,他們之間的關連仍然是他們的作者特徵向量之間夾角的餘弦值,但如果他們曾經合著過,在計算他們之間的關連時,先將他們合著論文的特徵向量排除在他們的作者特徵向量之外,在進行餘弦值計算,所以在計算每位作者和其他作者之間關連時所使用的作者特徵向量可能是變動的,因此稱為動態。基礎的主題模型假設每一個論文都是主題的混合(mixture),而每一個主題則都是詞語的混合。對於每一個論文,它的主題混合由一個已知參數α的Dirichlet分布所產生;每一個主題的詞語混合則由另一個已知參數β的Dirichlet分布所產生。在產生論文d前先根據Dirichlet分布Dir(α)取樣產生它的主題混合θd,然後再產生這個論文裡的每一個詞語,每一個詞語的產生是根據從主題混合θd中取樣得到的主題z以及其相對應的詞語混合ϕk所產生。本研究採用Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010)將作者資訊加入而擴充的LDA模型-- 作者-主題模型(author-topic model),這個模型假定每個作者是由一個已知參數α的Dirichlet分布所產生的主題混合。假設一個論文的作者群為ad ,在產生這個論文的每一個詞語時,首先從ad 中隨機抽取一個作者x以及他的主題混合θx,然後其主題z便由θx取樣產生。本研究利用Gibbs取樣(Gibbs sampling, Griffiths & Steyvers, 2004)進行作者-主題模型推論,產生包含每一個主題在詞語上的分布情形以及對每一位作者產生他在各主題上的分布情形等結果。因此利用作者-主題模型可以根據他們在主題分布的相似度測量他們的關連。

本研究的資料範圍為2000到2010年出版的圖書資訊學相關的八種主要期刊的 5227筆書目紀錄,從其中的 6282位不同的作者內選取50位最多產的作者。利用靜態特徵、動態特徵、主題模型和共被引分析等四種方法測量多產作者之間的關連並利用MDS (multidimensional scaling)和階層式叢集分析(hierarchical cluster analysis)進行視覺化。本研究在利用主題模型測量作者之間的關連時使用以下的參數,α設為50/K,其中的K是主題的數量,本研究設為20,β設為0.01,Gibbs取樣的迭代(iteration)次數設為1000次。針對每一對作者的四種關連測量方法所得到的值進行相關分析(correlation analysis),結果發現靜態特徵與動態特徵之間有最高的相關值,主題模型和其他兩種以內容為基礎的測量方法的相關值也較共被引方法來得高。四種測量方法皆可以發現LIS領域的兩大主軸:一個主軸是資訊檢索(information retrieval)與網路研究(web studies),另一則是科學評鑑(scientific evaluation)的測量指標(metrics)研究,LDA模型則在階層式叢集分析上有最連貫的結果。另外,以內容為基礎的方法比以引用為基礎的方法更容易解釋產生的結果。

In this study we present static and dynamic word-based approaches using vector space modeling, as well as a topic-based approach based on latent Dirichlet allocation for mapping author research relatedness.

Outcomes for the two word-based approaches and a topic-based approach for 50 prolific authors in library and information science are compared with more traditional author cocitation analysis using multidimensional scaling and hierarchical cluster analysis.

Science mapping is one of the most useful tools to visualize scientific structure. It helps to identify scientific themes, and discover new knowledge.

The unit of interest for mapping may include authors, articles, and journals.

To date, five approaches have been used to measure the relatedness between authors, where the nature of the relationship studied is based on the data used: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling analysis, and co-word analysis.

Recently, more sophisticated hybrid methods (i.e., using textual content and citations) have been applied to the mapping of articles (Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005) and journals (Liu et al., 2010).

As an initial investigation of these topics, our focus will be on authors whose publications appear in the highest impact library and information science journals.

In reviewing visualization studies for knowledge domains, Börner, Chen, and Boyack (2005) categorized relatedness measures into two broad categories: citation linkages and co-occurrence similarities.Within the relatedness measures, five basic approaches were identified: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling, and co-word analysis.

Direct citation accounts for the relatedness between a citing work and a cited work based on citing behavior. ... Shibata, Kajikawa, Takeda, and Matsushima (2008) explored citation networks for two research domains and divided the networks into clusters in order to identify research fronts. Direct citation has not attracted wide attention. One possible reason may be its requirement for a very long time window to obtain a sufficient linking signal for clustering (Boyack & Klavans, 2010).

The idea that two articles that share the same references are related, referred to as bibliographic coupling, was outlined by Kessler (1963). The more references two articles have in common, the more closely related they are thought to be. Note that this list is static over time because references within articles do not change. With the interrelation of this link, scientific products can be ordered into groups. Weinberg (1974) reviewed the theory and practical applications of bibliographic coupling and granted the usefulness of the method. More recently, Zhao and Strotmann (2008) aggregated bibliographic coupling at an author’s oeuvre (body of work) level, which they called author bibliographic-coupling analysis (ABCA). They found ABCA can provide an effective picture of current active research in a field.

Cocitation analysis, introduced by Small (1973), is probably the most influential approach for assessing relatedness measures. If two articles are cited by the same third article, these two articles are co-cited. The assumption is that the appearance of two articles in the same reference list indicates a semantic association between the articles. Unlike traditional bibliographic coupling, cocitation is a dynamic relationship based on the citing authors. New citing authors can change the cocitation relationship. This feature is important because science is developing continuously. Relationships among scientific units being studied should be able to incorporate this dynamic change.

White and Griffith (1981) first applied cocitation techniques to authors, called author cocitation analysis or ACA. The essential transformation is to consider “Author” as a body of writings by a person (i.e., an oeuvre). So the cocitation of authors applies to any work by any author being co-cited with any work by another author.

Since then, a number of studies have been conducted using variations of the ACA method, including normalization (Ahlgren, Jarneving,&Rousseau, 2003; Leydesdorff&Vaughan, 2006; White, 2003; van Eck & Waltman, 2009), author counts (Zhao & Strotmann, 2011), and last-author ACA (Zhao & Strotmann, 2010).

One disadvantage of cocitation analysis is the lack of cognitive interpretation of the relatedness of the co-cited units. Without enough domain knowledge, one can hardly interpret the cocitation map.

Leydesdorff (1987) argued that cocitation maps only partially represent the structure of science.

A co-authorship relationship is established when authors co-publish a paper. Glänzel (2001) studied international co-authorship links to reveal the structures in international collaborations. Liu, Bollen, Nelson, and Van de Sompel (2005) constructed a network with co-authorship relations in the field of digital libraries. Ding (2011b) studied scientific collaborations and citation patterns of researchers and combined the results with a topic model approach to examine collaborations among researchers who share similar and different research interests.

It is this feature of co-authorship that makes co-authorship analysis more revealing of a social network rather than a scientific structure.

Co-word analysis collects evidence of relatedness from co-occurring keywords from different articles. Compared with the approaches introduced earlier, co-word analysis directly uses actual contents to measure relatedness, whereas the others find indirect evidence through citation and co-author relations. An obvious advantage of co-word analysis is that relatedness can be interpreted directly according to document contents.

Coulter, Monarch, and Konda (1998) mapped the discipline of software engineering with co-word analysis. Indexing terms from the ACM Computing Classification System were used as the unit of analysis. Ding, Chowdhury, and Foo (2001) conducted a co-word analysis on a sample of 2,012 articles from the Web of Science (WoS) to reveal themes of information retrieval research.

Leydesdorff (1997) noted that the meaning of words change from position to position and from one text to another. He also suggested this change will destabilize the science map produced by co-word analysis.

Another disadvantage of using indexer-assigned keywords as the source for co-word analysis is the “indexer effect” (Law & Whittaker, 1992), which creates bias through factors such as the artificiality of an indexing language, delays in changes to the indexing language to reflect the current state of a discipline, and subjectivity in the assignment of index terms.

In the vector space, a number of documents constitute a document space. The centroid of the document space is a summarization of the characteristics of the space. It represents the average vector for a group of documents.

Each author will be viewed as a document space consisting of the articles he/she has written. This space is a subspace of the collection space, named the author space. The centroid of the author space will be used to represent the author. The relatedness between authors will be measured through the similarity between the centroids of their author spaces.

The topic model is an improvement over the basic vector space model in terms of relieving the independence assumption and capturing the term associations. Instead of assuming independence among terms, the topic model assumes exchangeability among terms in documents, which is a much looser assumption.

Early works on the topic model include latent semantic indexing (LSI) by Deerwester et al. (1990) and the probabilistic LSI (pLSI) by Hofmann (1999). LDA is a more recent technique proposed by Blei, Ng, and Jordan (2003). It has an advantage over LSI in explicitly modeling the latent topics, and over pLSI in solving the overfitting problem (i.e., a model with too many parameters).

The LDA model treats a document as a mixture of topics and a topic as a mixture of terms. Each document (i.e., a mixture of topics θ) is generated from a latent Dirichlet distribution with a prior of α, and each topic (i.e., a mixture of terms ϕk) is generated from a Dirichlet distribution with a prior of β. The generation process entails, first, sampling a document θd from Dir(α). At each position of a word in a document, a topic z is selected according to θd, and a word w is selected according to z and ϕk.



Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010) extended the original LDA model to include authors and proposed the author-topic model (Figure 2). This model includes authorship information in the generative process. Each document has a number of authors ad. Each author is considered as a distribution of topics drawn from a Dirichlet distribution with a prior of α. For each word in a document, an author x is randomly drawn from ad and the topic distribution associated with this author is θx. Then a topic z is selected the same way as in a LDA model to generate the observed word w.



The advantage of this author-topic model is that it adds authorship information to the model, so that the topics are learned and assigned to documents accordingly. In the output of this model, each author is a distribution of different topics; each topic is a distribution of terms. As the purpose of the current study is to measure the relatedness of authors, the author-topic model will be appropriate to produce author similarities based on their topics.

Gibbs sampling (Griffiths & Steyvers, 2004) is used to estimate the parameters in the model.

Table 1 lists the eight journals selected for inclusion in the study. ... Bibliographic records for documents published in these journals between 2000 and 2010 were downloaded. Records downloaded were further limited to three document types: articles, proceedings papers, and reviews. ... In total, 5,227 records were downloaded from WoS. The raw WoS records were processed, and only three fields were kept: the article title (i.e., “TI” field), the Keywords Plus (i.e., “ID” field), and the abstract (i.e., “AB” field). The records then were indexed with the widely used Lemur information retrieval toolkit (http://www.lemurproject.org/). Stop words were removed and stemming was applied.



From the 5,227 records downloaded, we were able to identify 6,282 different author names using string matching. Because it is impractical to map all of the authors in our collection, we selected the 50 most prolific authors according to the WoS “analyze results” function. ... We selected the most prolific authors because the more an author writes, the better the algorithm used “understands” her/his interests, and thus the more accurate our assessment will be.

For each author in our author list we then generated an author space consisting of all the articles he/she wrote. TF*IDF term weighting was employed to assign term significance in the space. Terms that were single characters or only consisted of digits (e.g., “2001”) were filtered out. We believe that these terms add noise to the space rather than meaning. The relatedness between authors is measured through the cosine between the centroids of the author spaces.

One could argue that this creates a biased assessment of the strength of the relationship because there is an exact match for the text of the co-authored publications that creates a stronger bond than for two authors who have published in a common area but did not collaborate. On the other hand, the simple fact that the collaboration has resulted in one or more co-authored documents should be acknowledged as a strong tie between the authors.

In a static space, each author has her/his own space that consists of her/his articles. This space does not change when measuring author relatedness. ... The relatedness of authors will include the similarity arising from the strength of the co-authorships.

Conversely, in the dynamic author space, the author spaces depend on a pair of authors. Co-authored articles by the pair of authors are excluded. In this case, each author may have a different author space when measured with different authors.

The vector space model provides a number of readily available measures of relatedness. The most popular is the cosine measure, which measures the cosine of the angle formed by two vectors in the space. It basically measures the term weight distribution between two vectors. The more similar the distribution is, the higher the cosine value is expected to be.

Gibbs sampling (Griffiths & Steyvers, 2004) was used to estimate the parameters in the author-topic model. We set the number of iterations to 1,000. The hyperparameter α was set to 50/K where K is the number of topics and hyper β is set to 0.01.We tested different K, or number of topics, values and decided to report the results from K = 20 because it produced the most reasonable outcome by our judgment.

The topic model toolbox was employed to perform the learning process (http://psiexp.ss.uci.edu/research/programs_data/toolbox.htm).

An author-topic LDA model (Rosen-Zvi et al., 2010) was trained on our collection and a pair-wise cosine similarity measure comparison of the 50 authors was conducted, resulting in a symmetric matrix of similarity values based on the LDA modeling. Similarity matrices were also calculated for both the static and dynamic author spaces. Multidimensional scaling was used to visualize the relationships among the authors. ... Because the data represent a type of similarity measure, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006). To provide additional insights into the grouping of the authors, hierarchical cluster analysis (complete linkage method) was used in SPSS to superimpose groups of authors on the MDS maps to provide an additional means to assess the coherence in the resulting proximities between authors.

After tokenization of the field contents, 916,383 tokens, or individual words, were identified; the number of unique tokens, or distinct words, was 12,537. The average document length was 175.32 tokens.

An examination of the pair-wise correlation of these author relatedness measures reveals significant and moderate level correlations between the word-based, topic-based, and author cocitation measures (Table 5). It is not surprising that the static author map has a high correlation with the dynamic author map (Kendall’s tau b = 0.971). Similarly, the correlations among the three content-based approaches are generally higher than their correlations with the cocitation approach. This provides preliminary evidence that they measure different types of relationships.

In all cases, the largest singular group consists of authors who work with different aspects of metrics-based studies, which is labeled as “Informetrics” in general in the two word-based maps and “Scientific impact evaluation” in the other two maps. This labeling indicates that the metrics-related topics have been a frequently investigated theme by the prolific authors in the selected journals during the first decade of the 21st century.

It is also noteworthy that the topic groupings of each of the maps largely aligns along the horizontal or vertical axis, with one side representing information retrieval (system and behavior) and web studies, with the other side corresponding to metrics-based or scientific evaluation studies.

As is shown from the maps, the static map (Figure 3) and dynamic map (Figure 4) are generally consistent in terms of the location of the authors, which indicates that the exclusion of similarities resulting from collaborations does not affect the overall layout. However, drastic changes may happen to individuals who have collaborated frequently with another author.

At the four-cluster agglomeration, the LDA map (Figure 6) provides the most coherent representation of the author map in relation to the generated clusters. At the two-cluster agglomeration, the clusters are neatly divided along the vertical axis, with metrics-related research represented on the left, and web and information retrieval-related themes on the right. Although the group membership of some individuals is still debatable, such as “Ingwersen_P” in the “Scientific impact evaluation” group given that he has also published in information retrieval and webometrics, the overall layout of the LDA map does provide semantically meaningful relationships.

Of the five author relatedness methods discussed earlier, only co-authorship provides a direct connection between authors.

Cocitations are contributed by third parties.

Direct citations reflect an author’s assessment of relatedness to a cited author or work but are still based on perception or the subjectivity inherent in citer motivation (Bornmann & Daniel, 2008).

This is also the case for bibliographic coupling, where the strength of the relationship is assessed by the overlap of references selected by two authors.

Co-word or topic-based studies can be argued to be the least influenced by citing behavior because they rely solely on the words developed by the authors themselves.

The newly proposed content-based approaches overcome several limitations of the more traditional cocitation approach.

In addition to avoiding citer subjectivity inherent in citation-based data, the links between authors will be more interpretable compared with the cocitation maps. The top terms/topics will be identifiable to help interpret the links between authors.

The content-based methods do not require an author to be cited in order to be included in the map. As long as the author has some publication record, her/his relatedness with other authors can be identified. This provides the opportunity for researchers who have not been widely cited to be included in the author map.

Furthermore, cocitation analysis outcomes may be affected by limited numbers of citations that do not reflect the true strength of the relationship between authors. This can be seen when comparing the cocitation outcomes with the topic-based outcomes, where several authors with low citation counts, and therefore low cocitation counts, end up at the periphery of the map. For the LDA outcome, these authors are more centrally situated among authors with similar topic areas.

The word-based and topic-based methods can be considered an extension of co-word analysis, where words are used to determine the relatedness of authors.

In Healey, Rothman, and Hoch (1986), a paradox is introduced: if a map represents a field that is already known to experts, then it is useless because it does not reveal anything new; if the map deviates from the expectation of the experts, then its outcome is questionable.

This initial investigation, which compares prolific authors from LIS, demonstrates: (1) the potential for more topically meaningful outcomes from the new methods when compared to more traditional cocitation analysis; (2) the topicbased method using LDA for the data used in this study produces more distinctive clusters and reasonable results than the two word-based approaches.