顯示具有 LIS 標籤的文章。 顯示所有文章
顯示具有 LIS 標籤的文章。 顯示所有文章

2015年4月21日 星期二

Tseng, Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

Tseng,  Y.-H. and Tsay, M.-Y. (2013) Journal clustering of library and information science for subfield delineation using the bibliometric analysis toolkit: CATAR. Scientometrics, 95, 503-528. doi: 10.1007/s11192-013-0964-1.

近幾十年來,發展出許多科學計量分析技術,包括為了群集(clustering)書目資料所需的各種相似度(similarity)計算技術,如共被引(co-citation)、書目耦合(bibliographic coupling)與詞語共現分析(co-word analysis),這些技術的比較分析可參見Yan and Ding (2012)。並且有很多可以在網路上自由下載使用的軟體工具製作並包裝這些技術,提供科學計量分析應用,知名的軟體工具如CiteSpace (Chen 2006, Chen et al. 2010)、Sci2 Tool (Sci2 Team 2009)、VOSviewer (Van Eck and Waltman 2010)、BibExcel (Persson 2009)及Sitkis (Schildt and Mattsson 2006),這部分的分析則可參見Cobo et al. (2011)。本研究包含兩個部分:提出包含一系列利用書目計量資訊進行群集與映射(mapping)技術的科學計量分析軟體工具集 CATAR,並且將此工具集應用於圖書資訊學(library and information science, LIS)領域後,希望能夠利用期刊群集的結果,確認與分析次領域,以及建議適合研究評估(research evaluation)用途的LIS期刊集合。

Åström (2002)從領域概念的視覺化研究獲得一個結論:期刊的選擇確實影響研究領域如何被知覺與定義,也就是研究領域的界定(delineation)與期刊的選擇有密切關係。已經有許多的研究對圖書資訊學進行次領域界定,而這些研究大多參考ISI的JCR主題分類中與圖書資訊學最相關的類別IS&LS(Information Science and Library Science)。IS&LS類別下並不只包含圖書資訊學的相關期刊,這個類別涵蓋兩個密切相關的領域資訊科學(Information Science)和圖書館學(Library Science),此一範圍與圖書資訊學有些微不同。根據Leydesdorff (2008),JCR主題分類以期刊的題名、引用模式(citation patterns)等等做為標準進行分類,但是這個分類結果與從資料庫本身的引用資料所產生的網路上的主要成分(principal components)得到的分類結果並不十分相符。因此次領域界定研究大多經過人為的挑選做為分析資料的期刊,並沒有完整收錄IS&LS主題下的所有期刊。

進行次領域界定時常使用的技術包括:利用共被引分析比較一對項目,利用凝聚式階層群集(agglomerative hierarchical clustering, AHC)將項目分群產生樹狀圖(dendrogram),利用多維尺度(multi-dimensional scaling, MDS)產生視覺化的二維或三維映射圖。若干重要的研究如:Åström (2002)從圖書資訊學重要期刊中選取1135篇出版在1998到2000年的文章,利用BibExcel軟體工具進行作者共被引(author co-citation)以及關鍵詞共現分析,並產生MDS映射圖,52位高被引作者的共被引產生三個群集:"硬"資訊檢索(hard information retrieval)、"軟"資訊檢索(soft information retrieval)以及書目計量學(bibliometrics),47個較常出現的關鍵詞則分為圖書館學(library science,LS)、資訊檢索(information retrieval,IR)及書目計量學。Åström (2002)認為作者共被引分析沒有出現圖書館學的原因可能與圖書館學研究的出版管道有關,如果引用的資料像是書籍或地區期刊沒有出現在JCR,圖書館學作者便無法出現在引用為基礎的排名上。Åström (2007)對55種在JCR 2003主題類別下的期刊,選擇21種圖書資訊學相關期刊的13605篇文章進行文件共被引分析,在從1990到2004年的三個時段發現圖書資訊學可分為資訊計量學(informetrics)和資訊搜尋與檢索(information seeking and retrieval)兩個穩定的次領域,而隨著全球資訊網的普及,網路計量學(webometrics)在兩個次領域上都成為主要的研究議題。Jassen et al. (2006) 對2002到2004年五種圖書資訊學相關期刊的938篇文章,應用一系列的全文分析技術以及MDS和AHC,將938篇文章分為六個群集:兩個群集與書目計量學有關、一個群集為IR、一個包含一般議題、另兩個較小但愈來愈重要的群集分別是網路計量學和專利分析(patent analysis)。Moya-Anegon et al. (2006)從24種較有影響力的期刊中選擇17種期刊,排除將資訊科學(information science, IS)應用到特定技術或知識領域(例如:醫學、地理學、電訊傳播等),從17種期刊引用的參考文獻,對77位最常被引用的作者和73篇最常被引用的期刊進行共被引分析,映射使用的技術包括MDS和AHC以及自組織映射圖(self-organizing map)。作者共被引分析的結果產生六個次領域:科學計量學、引用分析、書目計量學、"軟"(認知導向)資訊檢索、"硬"(演算法導向)資訊檢索以及傳播理論(communication theory)。而期刊共被引分析的結果則有四個群集:IS、LS、科學研究(science studies)以及管理學(management)。在期刊共被引分析的科學研究大致上可以對應為作者共被引分析的科學計量學、引用分析、書目計量學,IS為"軟"資訊檢索和"硬"資訊檢索。如Åström (2002)同樣的原因,LS沒在作者共被引分析的結果當中。Waltman et al. (2011)以JASIST為種子,選擇與該期刊共被引較多的期刊,連JASIST共48種,進行期刊的書目耦合(bibliographic coupling)分析,並且利用VOSviewer呈現視覺化結果,共分為LS、IS以及科學計量學等3個次領域。Milojevic et al. (2011)使用詞語共現分析探討1998到2007年出版的16種期刊上的10344篇文章,16種期刊根據Nisonger and Davis (2005) 的研究所挑選,分析100個文章題名上最常出現的詞語,進行共現分析,並以AHC歸類,結果三個主要群集為LS、IS以及書目計量學/科學計量學。

Åström (2002)以關鍵詞的共現分析所得到的結果包括LS次領域,但作者共被引分析所得到的映射圖上並沒有產生這個次領域。Moya-Anegon et al. (2006)的期刊共被引分析與作者共被引分析也略有不同,期刊共被引分析的結果上有作者共被引分析沒有的LS和管理學兩個次領域,反之,作者共被引分析的結果上則可以發現期刊共被引分析沒有的傳播學理論(communication theory)。一般認為這和作者引用的行為有關,LS作者的引用次數大多沒有達到分析的門檻,因此無法在上述兩個研究的作者共被引分析結果上呈現。

Ni et al. (2012)從JCR的IS&LS類別下的61種期刊,排除3種非英語的期刊,將選取的58種期刊進行場域-作者耦合(venue-author coupling)、期刊共被引分析、詞語共現分析、期刊連結(journal interlocking)等四種分析。分析的結果再進行MDS與AHC分析,四種方式所得到一致的次領域包括:管理資訊系統(managment information systems, MIS)、IS、LS和特殊化群集(specialized clusters),並且在四種方法所得到MDS映射的圖形上都可以發現MIS與其他群集分離,Ni and Ding (2010)與Ni and Sugimoto (2011)建議JCR上的圖書資訊相關期刊應進行適當的重組。

本研究(Tseng and Tsai 2013)應用的資料範圍為2000到2004與2005到2009在Web of Science 的Journal Citation Report中 Information Science & Library Science (IS&LS)主題分類下的所有期刊,在前期(2000~2004年)共50種,後期(2005~2009年)共66種。本研究的分析程序採用Borner et al. (2003)整理的一般工作流程,步驟包括:1) 資料蒐集(data collection);2)文本分段(text segmentation);3)相似性計算(similarity computation);4)多階段群集(multi-stage clustering);5)群集標名(clustering labeling);6)視覺化(visualization);7)面向分析(facet analysis)。這些步驟中所需的技術都已經整合到軟體工具CATAR(Content Analysis Toolkit for Academic Research, http://web.ntnu.edu.tw/~samtseng/CATAR/)上。在計算文件間的相關性時,本研究以一種期刊做為一個文件,所有論文引用的期刊做為文件的特徵,然後利用Dice係數(Salton 1989)計算期刊相似性,例如兩種期刊X與Y,R(X)與R(Y)分別是它們引用的期刊,它們之間的相似性計算為Sim(X, Y) = 2 ∙ |R(X)∩R(Y)|/(|R(X)|+R(Y)|)。也就是利用書目耦合計算期刊之間的相似性。期刊的群集則是利用完全連接階層群集法(complete-linkage hierarchical clustering)。首先將每個文件視為一個群集,然後將一對最相似的群集合併起來,產生一個較大的群集,然後重複進行上面的步驟,而兩個群集的相似性定義為兩個群集間最小的文件相似性,如果相似性超過某個預先設定的閾值,便將兩個群集合併,一直到無法再產生合併為止。此外,本研究採用Silhouette指標(Ahlgren and Jarneving 2008; Rousseuw 1987; Jassen et al. 2006)。

此一研究的資料包含JCR的IS&LS主題下的期刊,分為2000-2004年與2005-2009年兩個時期,前一個時期包含50種期刊,9546筆論文資料;後一時期則有66種期刊,11471筆論文資料。從群集結果的樹狀圖(dendrogram)和MDS映射的結果顯示,IS&LS主題下的期刊在兩個時期都有IR、MIS、科學計量學、學術圖書館(academic library)、醫學圖書館(medical library)、館藏發展(collection development),以及開放取用(open access)和地區圖書館(regional library)兩個後期出現並且較小的群集。並且MIS群集的期刊在知識基礎(intellectual base)上與IS&LS主題的其他期刊分離,表示這群集下的期刊具有較特殊的引用模式。本研究以期刊的書目耦合進行分析,從期刊知識基礎(intellectual base)得到MIS群集與其他分離的研究結果,與Ni et al. (2012)利用期刊共被引分析、期刊連結、術語使用(terminology usage)和合著(co-authorship)研究等不同方法的研究結果相同,這也為許多探討圖書資訊學認知結構的研究認為不應將MIS相關期刊與其他期刊包含在ISI的同一個主題IS&LS下,在進行分析時需要排除MIS相關期刊提供了佐證(Larivière et al. 2012)。此外,並且以多樣性指標(diversity index)分析群集特性,揭露出某些次領域具有地區(regional)特性。

2015年4月2日 星期四

Wang, F., & Wolfram, D. (2014). Assessment of journal similarity based on citing discipline analysis. Journal of the Association for Information Science and Technology.

Wang, F., & Wolfram, D. (2014). Assessment of journal similarity based on citing discipline analysis. Journal of the Association for Information Science and Technology.

利用Web of Science的主題分類,計算引用期刊的學科頻率分布能夠提供被引用期刊進行相似性比較的特徵,相較於共被引方法,這種相似性比較的維度較小,可以減少許多計算量。本研究比較Web of Science的資訊科學與圖書館學主題分類下的40種高影響力期刊,並以多維尺度法(multidimensional scaling)和階層式群集分析(hierarchical cluster analysis)比較比較所提出的方法與共被引方法的相似性估算結果。分析期刊的出版時間範圍為1987到2011,以5年為一個時期進行分析。在各期刊中,以Scientometrics (SCI)以及Journal of the Association for Information Science and Technology (JASIST)的引用期刊分布的學科最多元,因為JASIST有較廣的涵蓋範圍以及其他領域都對測量研究(metrics research)感到興趣。產生的映射圖與群集結果顯示某些期刊並不接近其他期刊。相似性估算結果顯示引用學科分析與共被引分析相似,各個時期兩種方法所得到的結果在分為三個群集的情況下,大多可以發現包含一個LIS群集、一個MIS群集以及一個較分散而邊緣的群集,不過組成群集的成員也有些不同,因此Wang and Wolfram (2014)建議可以引用學科分析做為共被引分析的補充。

The frequency distribution of disciplines by citing articles provides a signature for a cited journal that
permits it to be compared with other journals using similarity comparison techniques.

As an initial exploration, citing discipline data for 40 high-impact-factor journals assigned to the “information science and library science” category of the Web of Science were compared across 5 time periods. Similarity relationships were determined using multidimensional scaling and hierarchical cluster analysis to compare the outcomes produced by the proposed citing discipline and established cocitation methods.

The maps and clustering outcomes reveal that a number of journals in allied areas of the information science and library science category may not be very closely related to each other or may not be appropriately situated in the category studied.

The citing discipline similarity data resulted in similar outcomes with the cocitation data but with some notable differences. Because the citing discipline method relies on a citing perspective different from cocitations, it may provide a complementary way to compare journal similarity that is less labor intensive than cocitation analysis.

The application of visualization techniques to groups of bibliographic entities (publications, journals, or authors) provides a method for assessing the closeness of relationships among entities of interest. ... On a fundamental level, these investigations allow us to understand better the structure of disciplines based on the production of scholarship and how this changes over time (e.g., White & McCain, 1998). On a more specific level, findings can help to assess the impact of entities of interest or to situate disciplines or specializations within a larger context.

Leydesdorff and Cozzens (1993) studied how to delineate and attribute journals to specialties based on journal−journal citations and their changes over time. They demonstrated how the data could be used to construct macrojournals, consisting of aggregations of journals around a central journal.

Pudovkin and Garfield (2002) developed a journal relatedness factor based on citing and cited journals. The method was proposed to help identify thematically related journals.

Similarly, Glänzel and Schubert (2003) proposed the categorization of journals using a three-step process involving predefined categories, journal classification, and article classification for articles in journals with ambiguous subject assignments based on references.

Rafols and Leydesdorff (2009) compared the outcomes of two algorithms for the decomposition of large matrices against Web of Science (WoS) subject categories and Glänzel and Schubert’s categorization. The four methods resulted in similar map outcomes on a large scale.

Leydesdorff and Schank (2008) visualized and animated the disciplinary ties of three seed journals over time to demonstrate relationships among journals and their interdisciplinarity.

Boyack and Klavans (2010) compared results from cocitation analysis, bibliographic coupling, direct citation, and a hybrid approach for accuracy of outcomes in representing research fronts for a large corpus of biomedical literature. They noted that bibliographic coupling performed the best in representing the research fronts.

White (2000) proposed the use of citers to identify characteristics of a given author’s research such as an author’s citation identity, which consists of all the authors a given author cites. White also introduced the idea of citation image-makers, consisting of the authors who refer to a cited author. The citation image-makers approach may also be applied to journals, where citing authors constitute the citation image-makers of the journal.

Yan, Ding, Milojević, and Sugimoto (2012) explored community structures in IR research by combining topic modeling and community detection with IR literature to reveal the changing landscape of IR research.

To reduce the dimensionality of the similarity comparison, disciplinary identifiers for citing articles/journals may be used to reduce the number of comparisons that have to be made.

For the purposes of this study, WoS research areas are used. In this paper the research areas are referred to as disciplinary assignments.

This research is guided by several questions.
1. Does the frequency distribution of disciplines of citing journals permit comparison of journal similarities in a meaningful way?
2. Are the results of such a comparison similar or complementary to the better-established approach of cocitation analysis?
3. Do the similarities among journals within the same disciplinary categorization change over time as reflected in the changes in the frequency distribution of citing journal disciplines?
4. Can these similarities (or distances) provide decision support for whether journals should be grouped together in citation index services such as Thomson Reuters’ Journal Citation Reports?

Forty high-impact journals included in the Thomson Reuters’ 2011 Journal Citation Reports grouped in the category ISLS were selected for the study.  ... In addition to many of the journals rated highly in library and information science (LIS), as evidenced by a perception study of LIS deans and Association of Research Library directors conducted by Nisonger and Davis (2005), this category includes journals in allied areas such as management information systems (MIS), geographic information systems, and medical informatics.

Among the 20 highest-impact journals listed in the ISLS category, only 3 are included in the top 20 journals rated by LIS deans based on their familiarity with these journals. The majority of the remaining journals in the top 20 based on impact factor could be argued to be from allied areas given their additional classification in other WoS research areas and the lack of familiarity or resulting lower prestige as determined by LIS deans.

Citing article/journal data were collected from 1987 to 2011 and were divided into 5-year intervals.

For each journal, all articles, review articles, and conference proceeding articles were selected; all other publication types such as cited material were excluded. For each time period, the “create citation report” in the WoS was selected to identify all citing articles. The number associated with “citing articles” was then selected to retrieve the list of citing articles. The WoS “analyze results” feature was next selected for the list of citing articles. On the results analysis page, “research areas” were selected as the ranking field to provide the tabulated list of citing disciplines. The ranked list of citing disciplines was then copied into an MS Excel spreadsheet.

The list of research areas and their frequencies represent the journal’s citing discipline profile for each time period.

Salton’s cosine similarity measures were calculated for each pair of journals to produce a symmetric
matrix of journal similarity values ranging between 0 and 1 (Ahlgren, Jarneving, & Rousseau, 2003, 2004; Egghe & Leydesdorff, 2009; Leydesdorff, 2006, 2007) for each time period.

To provide a baseline comparison, a cocitation analysis was also conducted with the same journals.

Multidimensional scaling (MDS) PROXSCAL analysis and hierarchical cluster analysis in SPSS v.20 were applied to the symmetric similarity matrices.

The PROXSCAL algorithm was used instead of ALSCAL for the MDS procedure because it allows similarity or dissimilarity matrices to be used and has been shown to provide superior results for cocitation studies (Leydesdorff & Vaughan, 2006).

For hierarchical clustering, Ward’s method was used. Minkowski distance and squared Euclidean distance were each explored and produced the same outcomes at the three-cluster level. Clustering outcomes were superimposed onto the MDS maps.

Library Resources and Technical Services (LRTS) consistently attracted citations from the fewest discipline areas, indicating a narrower interdisciplinary focus. In fact, the number of citing article disciplines has declined over the past decade for this journal, possibly indicating even narrower interdisciplinary impact.

Scientometrics (SCI) and the Journal of the Association for Information Science and Technology (JASIST), on the other hand, at different time periods each attract the most disciplinarily diverse citations. These outcomes are not unexpected given the broad coverage of JASIST and the interest in metrics research by other disciplines.



The MDS map of the journals using the proposed citing discipline approach for the first period appears in Figure 1. Among the journals, 14 of the 22 are situated in close proximity. A secondary group with three journals is situated on the periphery.

In combination with the cluster-analysis groupings, one can see at the three-cluster level that the tightly clustered journals are core to LIS.

A more widely dispersed second cluster of five journals consists of LIS and allied area journals in MIS. ... It is interesting to note that Government Information Quarterly (GIQ), International Journal of Geographical Information Science (IJGIS), and Journal of the Medical Library Association (JMLA)—at the time, still the Bulletin of the Medical Library Association–are situated more closely to and are clustered with the journals associated with the MIS area.

A peripheral “Other” cluster contains three journals.  ... Telecommunication Policy (TP), Journal of Scholarly Communication (JSP), and Social Science Information (SSI) are situated on the periphery of the map for the first and second time periods, indicating little similarity with the other journals in the citing discipline distributions.


The equivalent cocitation analysis map (Figure 2) at the three-cluster level, produces similar outcomes, but with several notable differences.

The International Journal of Information Management (IJIM) is situated more closely to LIS journals than to those in MIS.

Two of the MIS journals are situated in their own cluster along with GIQ and TP, equivalent to the “other” category. IJGIS appears at the periphery of the map in the MIS category.

The remaining journals are subdivided into two clusters that may be characterized broadly as information science and library science, respectively, with JSP and SSI being a part of these clusters.

There is a 63.6% overlap (14 of 22 journals) in the cluster assignments, indicating that there is a moderate level of agreement between the two approaches.

For the second time period, the three clusters for the citing discipline-based analysis consisted of a group of 12 journals representing the LIS area, an emerging cluster of journals focusing on the MIS area and several journals in allied areas, and an “other” group consisting of JSP, SSI, TP, and IJGIS.

The cocitation analysis outcomes for the second time period reveal a similar mapping arrangement and clustering of journals, with 15 journals corresponding to the LIS category, eight representing a group with an MIS focus, and an “other” category consisting of journals in allied areas.

The citing discipline MDS map and cluster analysis results for the three-cluster level are similar to the first two time periods, but with more distinctive LIS, MIS, and other clusters as the number of journals in each cluster has grown.

The cocitation analysis map and resulting clusters at the three-cluster level consist of primarily LIS journals, those in MIS, and the other category similar in composition to the citing discipline outcome. ... The cluster assignment match at the three-cluster level between the citing discipline and cocitation analysis methods is 88% (29 of 33 journals), indicating a high level of agreement.

The citing discipline MDS map for the fourth time period is similar to that for the previous time period.

Of note with the cocitation cluster analysis outcome for the fourth time period is a much larger other category that includes a number of journals categorized as LIS by the citing discipline method. GIQ and INFSOC are situated between the LIS and MIS groups, although they are placed in the other group.

The citing discipline and cocitation maps for the fifth time period appear in Figures 5 and 6, respectively. The outcomes for the citing discipline approach are quite similar to those for the third and fourth time periods, with well-defined LIS and MIS categories and a more scattered other category on the periphery.

The three clusters based on the cocitation analysis data again reflect the LIS, MIS, and other groupings. There are fewer members in the other category than for the fourth time period

Much in the same way that dimensionality reduction used in certain statistical methods and IR allows for simplified comparisons, the use of the WoS research areas by citing journals and their frequency instead of citing authors or citing journals provides a less computationally intensive way to assess journal similarity by reducing the dimensionality of the comparisons and the computational overhead.


Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

揭露與更好地了解科學傳播 (scholarly communication) 的中研究人員、研究團隊、機構、地區/國家、學科、出版品之間的關係,可以在不同尺度下進行研究。分析資料之間的連結可以從直接引用、共被引、共同著作、詞語或主題的共現,或者隱含式主題等形式進行。期刊間的相似性經常利用共被引,過去曾經進行期刊共被引研究的學科有經濟學 (McCain, 1991)、資訊檢索 (Ding, Chowdhury & Foo, 2000)、資訊系統 (Marion, Wilson & Davis, 2005)、醫療資訊學 (medical informatics, Morris & McCain, 1998)、類神經網路 (neural networks, McCain, 1998)以及半導體研究 (Tsay, Xu & Wu, 2003)。然而共被引研究有以下的困難:首先,在多個學科裡的應用,期刊與期刊間的共被引矩陣可能相當稀疏(Boyack, Klavans, & Börner, 2005)。其次,若是沒有Web of Science等資料庫來源,共被引資料的取得將會相當困難。最後,共被引分析大多只利用共被引次數,而沒有考慮引用來源的任何特性。

資訊計量學的重要研究之一是利用引用資料來確認期刊的學科或專業背景,錯誤的分類結果將會影響期刊在領域內的排名。Glänzel and Schubert (2003)發展一個三個步驟的期刊分類程序,使得期刊裡的文章可以根據參考文獻指定主題。Rafols and Leydesdorff (2009)對Web of Science的主題分類(Subject Categories)和Glänzel and Schubert (2003)的主題分類,比較兩種大型矩陣的分解演算法。Leydesdorff and Rafols (2009) 也使用主題分類引用頻率的引用矩陣研究170多種Web of Science的主題分類之間的關係。Leydesdorff and Schank (2008)以視覺化及動畫的方式呈現期刊之間的關係與它們的跨領域性。

過去的研究有利用期刊引用形象(journal citation image),也就是引用目標期刊的所有期刊的列表,做為期刊之間相似度評估的特徵。在計算上,可將期刊引用形象加入引用的頻率分布做為目標期刊的一種特徵(signature)。然而由於具有影響力與聲譽的期刊可能有相當大量的引用期刊,造成期刊引用形象的計算量較大。Wang and Wolfram (2014)提出使用引用期刊所屬的學科,利用引用學科的引用頻率做為期刊的特徵,來降低計算量。Wang and Wolfram (2014),提出引用學科分析(citing discipline analysis)來評估被引用期刊(cited journals)之間的相似性。引用學科分析根據目標期刊的引用期刊(citing journals)在Web of Science的研究領域上的頻率分布,取代直接。Wang and Wolfram (2014)並將引用學科分析應用於JCR的資訊科學與圖書館學 (Information Science & Library Science)的40種期刊,他們發現在多元尺度與群集分的結果中,若干期刊與其他期刊並不接近,有些同時被歸類於其他學科的期刊並不接近於資訊科學與圖書館學的期刊,而且當這些期刊同時也歸類於較大的相關領域時,其具有較高的影響係數(impact factors)將會降低其他期刊的排名。由於Wang and Wolfram (2014)只探討一個學科內的期刊,無法了解多個學科的期刊是否也具有同樣的情形。

本研究同樣利用引用學科分析估計期刊間的相似性。本研究使用來自於6個學科的120種期刊做為研究資料,其中5個學科彼此間較為接近,包括傳播學 (Communication)、電腦科學-資訊系統 (Computer Science-Information Systems)、教育學與教育研究 (Education & Educational Research)、資訊科學與圖書館學 (Information Science & Library Science)、管理學 (Management),另一個學科-地理學 (Geology) 則較遠。選取期刊的出版期間為1987到2012年,分為三個時期1987–1995、 1996–2004和 2005–2012。利用餘弦測量計算期刊間的相似性,並將估計結果應用於多元尺度(multidimensional scaling)、階層式群集分析(hierarchical cluster analysis)、主成分分析(Principal Component Analysis)等技術。

第一時期可以發現六個學科的相關期刊分為五個群集,其中電腦科學-資訊系統的期刊因為在這時期的數量較少而與資訊科學與圖書館學形成一個群集,地理學的群集與其他群集的距離相當遠。原本主題被分類在傳播學的 Journal of Advertising Research(JAR),在這時期的結果裡與管理學的相關期刊較接近。MIS相關期刊以及 Telecommunications Policy (TP)在資訊科學與圖書館學與管理學之間,Social Science Computer Review (SSCR)則是在資訊科學與圖書館學、傳播學與管理學三個學科間,另外Science Communication (SCOMM)雖然主題被分類在傳播學,但在本研究裡的結果,則是歸入資訊科學與圖書館學的群集。


第二時期電腦科學-資訊系統已經和與資訊科學與圖書館學分開,120種期刊共形成6個群集。在這時期的結果,部分主題分類於資訊科學與圖書館學的期刊被歸入其他學科,包括The Journal of the American Medical Informatics Association (JAMIA)和The International Journal of Geographical Information Science (IJGIS)被歸入電腦科學-資訊系統,後者甚至也很接近地理學;Decision Support Systems (DSS)則是原本在電腦科學-資訊系統分類下,而被歸入資訊科學與圖書館學的群集。Journal of Health Communication (JHC)的主題分類有傳播學和資訊科學與圖書館學,但在這時期更靠近傳播學的相關期刊而被歸類在傳播學的群集內。此外,原本分別在教育學與教育研究以及傳播學的the Academy of Management Learning& Education (AMLE)和JAR都被歸類在管理學的群集裡。

第三時期的六個群集對應到六個學科,但原先主題分類為傳播學的一些期刊更靠近於管理學期刊,而同時具有教育學與教育研究以及資訊科學與圖書館學兩種主題的International Journal of Computer-supported Collaborative Learning (IJCSCL),在結果上明顯可歸類為教育學與教育研究的群集。

有些被指定在某一個學科的期刊結果更接近其他學科,但有些被指定在多個學科的期刊則發現只有接近其中的一個學科。因此這個研究所提出的方法在計算期刊間的接近程度時能夠與期刊共被引分析(journal co-citation analysis)等傳統方法互補。由於以下的幾種原因,有愈來愈多的期刊跨越學科的邊界:1) 期刊出版的範圍愈來愈跨學科(more interdisciplinary),因此吸引愈來愈多其他學科的出版品引用。2) 整合式搜尋工具愈來愈普遍,更容易讓其他學科的作者發現。3) 愈多的期刊加入分析,使得期刊在引用學科的特徵上獨特性減少,彼此間更加相似。

為了更加了解跨學科的期刊,本研究利用主成分分析探討期刊的引用學科特徵。在第一時期,資訊科學與圖書館學的相關期刊中,共有五種期刊屬於兩種成分,包含ISJ, ISR, MISQ等三種管理學期刊以及屬於傳播學和電腦科學-資訊系統的SCOMM和DSS。但在資訊科學與圖書館學主題分類下的ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS以及 MISQ同時也被分類在電腦科學-資訊系統主題內,但卻沒有出現在電腦科學-資訊系統的成份裡。並且JAR和 AMLE分別被分類在傳播學和教育學與教育研究,但在本研究的第一時期結果只屬於管理學。

在第二時期的結果中,IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ等許多MIS期刊同時出現在管理學和資訊科學與圖書館學的成份裡。雖然SSCR僅有被分類在資訊科學與圖書館學主題,但同時包含在資訊科學與圖書館學與傳播學的成分中。另外,雖然IJGIS被分類在資訊科學與圖書館學主題,但並沒有出現在任何成分,顯然是六個學科較周邊的期刊。

第三時期同時在管理學和資訊科學與圖書館學成份裡的MIS期刊更增加了。沒有出現在任何成分的期刊,除了IJGIS以外,還增加了Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM)。

A similarity comparison is made between 120 journals from five allied Web of Science disciplines (Communication, Computer Science-Information Systems, Education & Educational Research, Information Science & Library Science, Management) and a more distant discipline (Geology) across three time periods using a novel method called citing discipline analysis that relies on the frequency distribution of Web of Science Research Areas for citing articles.

Similarities among journals are evaluated using multidimensional scaling with hierarchical cluster analysis and Principal Component Analysis.

The resulting visualizations and groupings reveal clusters that align with the discipline assignments for the journals for four of the six disciplines, but also greater overlaps among some journals for two of the disciplines or categorizations that do not necessarily align with their assigned disciplines.

Some journals categorized into a single given discipline were found to be more closely aligned with other disciplines and some journals assigned to multiple disciplines more closely aligned with only one of the assigned disciplines.

The proposed method offers a complementary way to more traditional methods such as journal co-citation analysis to compare journal similarity using data that are readily available through Web of Science.

Aspects of scholarly communication may be investigated from different levels of granularity to reveal and better understand relationships between researchers, research groups, institutions, regions/nations, specializations/disciplines, publications or publication outlets.

Connections that exist between sources of interest may take the form of direct citations, co-citations, co-authorship, co-occurrence of words or subjects, or more recently, latent topics.

Journal similarity comparison has been frequently studied using co-citations. Journal co-citation studies have been carried out on a number of fields including economics (McCain, 1991), information retrieval (Ding, Chowdhury & Foo, 2000), information systems (Marion, Wilson & Davis, 2005), medical informatics (Morris & McCain, 1998), neural networks (McCain, 1998), and semiconductor research (Tsay, Xu & Wu, 2003). 

One reason co-citation studies tend to focus on individual fields is that the journal–journal co-citation matrix that emerges when multiple disciplines are employed can be quite sparse (Boyack, Klavans, & Börner, 2005).

Co-citation data can also be labor-intensive to extract and are not easily available through citation database sources such Thomson Reuters Web of Science (WoS) without downloading all references from a corpus of articles.

Citation-based data may also be used to identify disciplinary or specialization affiliations for journals. This is particularly important for informetrics studies, where the misclassification of journals may affect the ranking of journals within a given field.

Pudovkin and Garfield (2002) developed a journal relatedness factor based on citing and cited journals. The goal of their proposed method was to help identify thematically related journals.

Similarly, Glänzel and Schubert (2003) developed a three-step process for the categorization of journals that involved pre-defined categories, journal classification and article classification for articles in journals with ambiguous subject assignments based on references.

More recently, Rafols and Leydesdorff (2009) compared the outcomes of two algorithms for the decomposition of large matrices against Web of Science Subject Categories and Glänzel and Schubert’s categorization. The four methods they used resulted in similar map outcomes on a large scale. Leydesdorff and Rafols (2009) also investigated the relationships among 170+ Web of Science Subject Categories using a citation matrix consisting of the subject category citation frequencies. They concluded that a classification scheme could be developed using analytical arguments.

Similarly, Leydesdorff and Schank (2008) visualized and animated the disciplinary ties of three seed journals over time to demonstrate relationships among journals and their interdisciplinarity.

Co-citation analysis relies on citing articles to identify the strength of relationships between the units of interest, whether authors, papers or journals; however, it does not consider any attributes of the source of the citations – only that the citations or co-citations exist. Authors such as White (2001) and Ajiferuke, Lu, and Wolfram (2010) have called for a shift in the focus of citation-based research away from citation counts received by an author of interest to the origin of the citation and its characteristics to assess author impact from a different perspective.

This research investigates the use of data derived from citing journals to assess the similarity of cited journals.

The journal citation image of a target journal, which is determined by the list of journals that cite the target journal, provides an indicator of the reach of a journal. When combined with the frequencies of citation by the citing journals, the frequency distribution of citations provided by the citing journals creates a “signature” for each cited journal. These signatures may be compared using various analytical methods.

One possible challenge associated with using the citing journals themselves to create a signature for a cited journal is the potentially high number of citing journals that an influential and prolific journal might attract.

Wang and Wolfram (forthcoming) proposed a method to reduce the computational overhead associated with the citing journal data. Their method of citing discipline analysis uses the subjects/disciplines assigned to the citing journal and the resulting citation frequencies of the citing disciplines to constitute the cited journal’s signature.

Wang and Wolfram (forthcoming) employed citing discipline analysis to explore journal similarity among 40 high impact journals in Information Science and Library Science (ISLS) as classified in Journal Citation Reports (JCR). They found that some of the journals classified into the ISLS category did not map in close proximity to one another based on multidimensional scaling and cluster analysis. A number of the journals included were also classified into allied fields, but did not cluster or appear in close proximity to a number of journals only classified in ISLS.

The authors noted that how journals are classified can impact journal rankings within a given field, where journals from related, but larger, fields may have higher journal impact factors (IF), which can reduce the rank of journals that are directly in the field. They observed that many of the high impact journals were from allied areas to ISLS.

One limitation of their exploratory study was the focus on a single discipline. Could similar affinities or differences in journal similarity be revealed using citing discipline analysis with journals from multiple fields? Also, by looking at multiple fields, does this pull journals also classified into other disciplines further out of an assigned category when included?

The present research is guided by the following questions:
1. To what extent are high impact journals from allied disciplines similar to one another based on the discipline of the articles that cite a given journal?
2. Do journals classified into multiple disciplines more closely align with one discipline than another or serve as bridges between the disciplines to which they are mapped based on the citing discipline distribution?

The field of Information Science and Library Science was selected as the seed discipline based on its interdisciplinary nature and familiarity to the authors. The top 20 journals based on 2012 JCR impact factors were selected. Four additional allied WoS disciplines were also selected based on the co-classification of journals appearing in the top 20 ISLS list with other disciplines, and the affiliation of information science and library science academic units with other disciplinary units, which demonstrates another type of alliance.

The four JCR disciplines selected comprised:
◦ Communication (COMM) – based on the existence of schools of communication & information.
◦ Computer Science, Information Systems (CSIS) – based on a number of iSchools and journal overlap in JCR.
◦ Education & Educational Research (EDER) – based on a number of ISLS units affiliated with colleges/schools of education.
◦ Management (MGMT) – based on the overlap of journals, particularly in Management Information Systems (MIS).

A sixth, more intellectually distant, discipline, namely Geology (GEOL), was also included. Geology was selected based on the outcomes of the UCSD Map of Science (Börner et al., 2012), where Earth Sciences were mapped as distant from the Social Sciences. By including journals from a more distant discipline, the ability for the citing discipline method to distinguish between more closely aligned and distant disciplines could be tested, where the distinctiveness of allied disciplines may be less defined by including a more distant discipline in the analysis.

A total of 120 journals were studied over the time period 1987–2012. To allow for a comparison over time, the journals were subdivided into three time periods: 1987–1995, 1996–2004, and 2005–2012. 

The data collection method for determining the frequency distribution of citing disciplines used in Wang and Wolfram(forthcoming) was adopted for the present study.

The “Create Citation Report” option in WoS was selected to identify all citing articles. The number of Citing Articles was then selected to retrieve the list of citing articles. The WoS “Analyze Results” feature was next selected for the list of citing articles. On the Results Analysis page, “Research Areas” were selected as the ranking field to provide the tabulated list of citing disciplines.

Salton’s Cosine measure was used determine the similarity between pairs of journals, resulting in a symmetric similarity matrix (Ahlgren, Jarneving, & Rousseau, 2003; Egghe & Leydesdorff, 2009; Leydesdorff, 2006).

Multidimensional scaling (MDS) analysis and hierarchical cluster analysis using SPSS v.20 were employed to visualize and categorize the relationships among the journals for each time period.

To provide a complementary analysis of the hidden groups that may be present in the data, SPSS’s Factor Analysis using Principal Component extraction with varimax rotation was also applied to the data using routines.


Fig. 1 shows the MDS locus of 70 selected journals in the first time period (1987–1995). The raw stress value was 0.01294,and the stress-I was 0.11376.

Only five clusters are shown because a distinctive sixth cluster did not emerge for this time period.

At the five-cluster level of assignment, journals from COMM, EDER, and MGMT form coherent clusters, although based on the MDS map some journals in each field are more closely located to journals in an allied discipline.

The fourth cluster combines journals from ISLS and CSIS. It is possible that the relatively small number of purely CSIS journals for this time period did not provide enough data for these journals to cluster into separate groups.

The MIS journals are situated in the ISLS cluster, but are located between the Library and Information Science (LIS) journals and MGMT journals.

The fifth cluster on the right side of the map consists of GEOL journals and, as would be expected, is quite distinctive from the other five disciplines.

The location of some journals on the map suggests that they are more similar to journals in one of the other given categories in JCR. The Journal of Advertising Research (JAR), for example, is situated with management journals but is only classified with COMM (and with Business, but this discipline is not included in this study).

Some journals classified in two disciplines served as a bridge between the two disciplines in the map. For instance, Science Communication (SCOMM), although classified with COMM journals, is situated in the ISLS cluster, but is in closer proximity to the COMM journals. In the remaining two time periods, SCOMM clusters with the COMM journals. Social Science Computer Review (SSCR), which is classified with ISLS and clusters with the discipline, is situated between ISLS, COMM and/or MGMT journals in each time period. The same is observed with Telecommunications Policy (TP), which bridges ISLS and MGMT for each of the periods of study.

The outcome for the second time period (1996–2004) appears in Fig. 2. The raw stress and stress-I values are 0.01648 and 0.12798, respectively.

In this map, 93 journals were categorized into six clusters, with CSIS separating from the ISLS cluster during this time period.

Of note is the greater number of journals assigned to one or more disciplines but aligning more closely to another discipline or only one of the assigned disciplines.

As an example, the Academy of Management Learning& Education (AMLE) and JAR are situated in the management cluster, and are located relatively far from their assigned disciplines, EDER and COMM, respectively.

Decision Support Systems (DSS) clusters with the ISLS journals but is classified in CSIS only. The same classification is observed for this journal in the third time period.

The Journal of the American Medical Informatics Association (JAMIA) is classified in ISLS and CSIS, but clusters with the CSIS journals for the remaining time periods.

Similarly, the Journal of Health Communication (JHC), which is classified in COMM and ISLS, does not appear to be similar to other ISLS journals and is situated more closely to COMM journals and clusters with them.

The International Journal of Geographical Information Science (IJGIS), which is classified with ISLS journals clusters with CSIS journals for this time period and the third period, although it is at the periphery of the cluster, perhaps indicating the CSIS discipline is the best match of the disciplines studied, but is not a very close match. Proximally, it is situated between CSIS and GEOL journals, which may indicate at least a peripheral similarity to some GEOL journals.

Again, the GEOL journals all cluster together farther from the other disciplinary groups.
Results for the third time period (2005–2012) appear in Fig. 3. The six clusters roughly correspond to the six disciplines. The results of raw stress calculation and the stress-I calculation are still relatively low, at 0.02224 and 0.14914, respectively.

Additional journals classified in COMM map closely to and cluster more closely with MGMT journals.

The International Journal of Computer-supported Collaborative Learning (IJCSCL) is classified with both EDER and ISLS but is clearly situated and clusters with the EDER journals.

Business Strategy and the Environment (BSE), although clustered with MGMT journals, appears to be pulled toward the GEOL journals, indicating a possible relationship with some of these journals.

Once again, the GEOL journals are distinctly clustered away from the remaining journals.

With each time period, more journals cross disciplinary boundaries by clustering with journals from allied disciplines or by mapping more closely to journals in allied disciplines. There may be several influencing factors to account for this observation.

First, the journals indeed may be becoming more interdisciplinary in their publication coverage, thereby attracting more citations from publications in other disciplines.

Second, the journals themselves may not be more interdisciplinary in their coverage, but are now more easily discovered by authors in other disciplines given the wider availability of federated search tools.

Third, with a greater number of journals included in the analysis for each time period, the distinctiveness of the citing discipline signatures may be decreasing, so some journals classified in allied disciplines may appear more similar to one another.

To determine common dimensions from the dataset, a Principal Component Analysis was conducted in SPSS for each time period. Outcomes for the Kaiser–Meyer–Olkin measure of sample adequacy (above 0.7) and Bartlett’s Test of Sphericity (p < .05) indicate the data were appropriate for PCA for all three time periods.


In this period, the six components explain 88.1% of the total variance that correspond to the six disciplines.

There are five journals underlined in the Table 2 that belong to two components, introducing inter-factorial complexity (Van den Besselaar & Heimeriks, 2001; Leydesdorff, 2007), including three MGMT journals (ISJ, ISR, MISQ).

SCOMM and DSS are assigned only to COMM and CSIS by WoS, respectively, but they also appear in the ISLS component, which supports the MDS and clustering outcomes. DSS continues to also load with ISLS for the remaining time periods.

Similarly, JAR and AMLE, journals classified by WoS only in COMM and EDER, respectively, load with the MGMT component for all the time periods in which they appear, but not their classified discipline, lending support for the re-classification of these journals.

The ISLS journals ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS, and MISQ are also classified at CSIS, but do not load with the CSIS component.

Outcomes for 1996–2004 appear in Table 3.

As with the first time period, several MIS journals (IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ) load to both MGMT and ISLS.

SSCR, which is classified with ISLS only, loads into the ISLS component, but also loads with a higher value into the COMM component, perhaps indicating the need for an additional classification assignment.

IJGIS does not load into any of the six components for second or third time period, lending evidence to the peripheral nature of the journal to the six fields studied and indicating it might be misclassified in ISLS.

Component outcomes for 2005–2012 appear in Table 4.

Similar to the previous time period, a growing number of MIS journals load to the ISLS and MGMT components.

As with IJGIS, two other journals, Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM) do not load into any of the six components, indicating a poor association with the six disciplines.

The boundaries between ISLS and CSIS are not as clear in the MDS and cluster analysis outcomes, where combinations of computer science, library and information science and management information systems journals may cluster together depending on the time period. These results may be influenced by the fact that a number of journals in the ISLS area are also categorized in the CSIS or MGMT category, thereby strengthening their relationships.

Despite the influence of the assigned discipline(s) – which then strengthens the disciplinary relationship(s) through journal self-citations, whether or not it is the best fit – some journals appear to be misclassified based on the disciplinary designations of the citations they attract.

As a prime example, JAR, which is only assigned to the COMM field, is situated in the MGMT category for each of the time periods studied for the MDS and clustering analysis as well as with the Principal Component Analysis.

A similar outcome is observed for AMLE for the two time periods in which it is included. It is classified as an EDER journal, but is situated with the MGMT journals based on the analyses conducted.

Several other journals are classified in more than one discipline, but clearly associate only with journals in one of the disciplines. JHC is classified in COMM and ISLS but clusters only with COMM journals. The same is observed for IJCSCL, which is classified in ISLS and EDER but groups only with EDER journals for each of the grouping methods used.

Other journals appear to move between disciplines over time. SCOMM is classified in COMM, but the MDS and clustering outcome places it initially with the ISLS journals for the first time period, and then in the COMM cluster and further away from ISLS in last two time periods.

Some journals, such as JCMC and SSCR, appear to be situated near borders between disciplines, which point to their interdisciplinary appeal or may indicate they serve as bridges between the disciplines.

A number of the ISLS journals that are considered to be library and information science journals (Nisonger & Davis, 2005) appear between CSIS journals and those in the MGMT cluster. In particular, MIS journals appear between MGMT and the ISLS/CSIS cluster for the first time period.

One application of citing discipline analysis that emerges from this analysis is that of decision support for the additional assignment or reassignment of journals to one or more disciplines.

Journal disciplinary classifications should be revisited over time to accommodate shifts in how journals are being cited by other disciplines. ... However, the shifts are at least an indication that the subject affiliations of the citing the journals are changing.

Citing discipline analysis provides, with some modest programming, a relatively easily implemented method for assessing the similarity of journals within disciplines or across allied disciplines that is computationally less expensive than using citing journal-based data.

The analyses reveal distinct groupings of journals based on their disciplinary assignments. As observed earlier in Wang and Wolfram (forthcoming), who only examined journals in a single field, the current research has demonstrated that citing discipline analysis can provide coherent and meaningful disciplinary groupings for journals in allied fields, even when journals from a more intellectually distant field are included.

The clustering and proximity of some journals classified in allied fields has changed over time, perhaps indicating a changing citing relationship between these fields.

2015年3月30日 星期一

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

確認科學領域的專業(specialties)本質是資訊科學的一項基本挑戰 (Morris & Van der Veer Martens, 2008; Tabah, 1999) 。由於1)可取用的書目資料來源愈來愈普及;2)網路上愈來愈多可提供分析與視覺化的電腦軟體工具;3)從多元來源而大量的資料吸收的要求愈來愈劇烈等原因,因此有愈來愈多的相關研究。共被引分析是對科學進行量化分析最常用的方法之一,特別是作者共被引分析 (author cocitation analysis, ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b)以及文件共被引分析 (document cocitation analysis, DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985)。作者共被引分析的目的在透過被相關文獻一起引用的作者群集,確認領域裡的專業。重要的作者共被引分析研究包括White & McCain (1998),這個研究以1972到1995年間12種資訊科學相關期刊的120位高被引作者進行作者共被引分析,研究結果發現當時的資訊科學分為兩個基本上彼此獨立的陣營:資訊檢索(information retrieval)與文獻(literature)。Zhao and Strotmann (2008a, 2008b) 以1996-2005年的資訊科學相關期刊資料重新進行了相同的研究,他們的結果發現了5個主要的專業:使用者研究(user studies)、引用分析(citation analysis)、實驗型檢索(experimental retrieval)、網路計量學 (Webometrics)以及知識領域的視覺化(visualization of knowledge domains),其中新興的兩個專業:網路計量學和知識領域的視覺化連繫了引用分析以及實驗型檢索,而使用者研究則是此時最大的專業。Aström (2007) 則是使用文件共被引分析的例子,他們分析了1990到2004年的21種圖書資訊學期刊,利用多維尺度法(multidimensional scaling, MDS)產生結果,他們的結果與White & McCain (1998)的研究類似,整個領域可分為兩個陣營,不過Aström (2007)的結果將稱為資訊尋求與檢索(information seeking and retrieval),而不是資訊檢索。

不管是作者共被引分析或是文件共被引分析其步驟大致如下:
1) 檢索引用資料。
2) 建構參考文件或作者共同被引用的矩陣。
3) 將共被引矩陣表示成節點與連結的圖(node-and-link graph)或是多維尺度法的組態(configuration),並且可以利用尋路網路(Pathfinder network scaling)或最小生成樹(minimum spanning tree)裁減連結。
4) 利用群集、社群發現(community finding)、因素分析(factor analysis)、主成分分析(principle component analysis)或者隱含語意索引(latent semantic indexing)等各種演算法確認專業。例如Morris & Van der Veer Martens (2008)、 Persson (1994)、 Tabah (1999)、 White & Griffith (1982)以及Janssens, Leta, Glänzel, and De Moor (2006)。
5) 根據群集成員間共同的主題(themes),解釋共被引群集的性質。通常需要豐富的領域知識,而且是一個花費大量時間與認知需求(cognitively demanding)的工作。

本研究對於作者共被引以及文件共被引形成的群集進行結構與動態的描述與解釋,分析的資料為1996到2008年間的12種資訊科學(information science)領域相關期刊,共計10853筆書目紀錄,引用的參考文獻為129060筆,引用次數為206180,而參考文獻的作者共有58711位。本研究以餘弦(cosine)測量作者或文件之間的關連大小,做為節點間的連結,建立網路;然後計算從原先網路導出的Laplacian矩陣(Laplacian matrices)的特徵向量(eigenvectors)找出群集。這種利用標準線性代數的頻譜群集(spectral cluster)演算法,較其他的群集演算法更有效率,而且因為不需要假設群集的形式,所以更有彈性與強健。標註群集方面則是利用引用文獻論文的詞語與摘要句,詞語包括題名與摘要中出現的名詞片語與索引詞(index terms),利用 tf*idf (Salton, Yang, & Wong, 1975)、對數似然比(log-likelihood ratio, LLR)測試 (Dunning, 1993)以及相互資訊(mutual information, MI)等三種資訊做為判斷的參考。摘要句則是從題名與摘要尋找最有代表性的句子,例如以Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007)對句子進行排序。




A multiple-perspective cocitation analysis method is introduced for characterizing and interpreting the structure and dynamics of cocitation clusters.

The generic method is applied to a three-part analysis of the field of information science as defined by 12 journals published between 1996 and 2008: (a) a comparative author cocitation analysis (ACA), (b) a progressive ACA of a time series of cocitation networks, and (c) a progressive document cocitation analysis (DCA).

Identifying the nature of specialties in a scientific field is a fundamental challenge for information science (Morris & Van der Veer Martens, 2008; Tabah, 1999).

The growing interest in mapping and visualizing the structure and dynamics of specialties is because of a number of reasons:
1. Widely accessible bibliographic data sources such as the Web of Science, Scopus, and Google Scholar (Bar-Ilan, 2008; Meho & Yang,2007) as well as domain-specific repositories such as ADS (http://www.adsabs.harvard.edu/) and arXiv (http://arxiv.org/).
2. Freely available computer programs and Web-based general-purpose visualization and analysis tools such as ManyEyes (http://manyeyes.alphaworks.ibm.com/) and Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/; Batagelj & Mrvar, 1998), special-purpose citation analysis tools such as CiteSpace (http://cluster.cis.drexel.edu/&u0007E;cchen/citespace/; Chen, 2004; Chen, 2006), and social network analysis such as UCINET (http://www.analytictech.com/ucinet6/ucinet.htm).
3. Intensified challenges for digesting the vast volume of data from multiple sources (e.g., e-Science, Digging into Data (http://www.diggingintodata.org/), cyber-enabled discovery, SciSIP; Lane, 2009).

Cocitation studies are among the most commonly used methods in quantitative studies of science, especially including author cocitation analysis (ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b) and document cocitation analysis (DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985).

For instance, once cocitation clusters are identified, assigning the most meaningful labels for these clusters is currently a challenging task because any representative labels of clusters must characterize not only what clusters appear to represent, but also the salient and unique reasons for their formation.

The new procedure reduces analysts' cognitive burden by automatically characterizing the nature of a cocitation cluster in terms of (a) salient noun phrases extracted from titles, abstracts, and index terms of citing articles and (b) representative sentences as summarizations of clusters.

ACA aims to identify underlying specialties in a field in terms of groups of authors who were cited together in relevant literature.

White & McCain (1998) presented a comprehensive view of information science based on 12 journals in library and information science across a 24-year span (1972–1995). It analyzed cocitation patterns of 120 most-cited authors with factor analysis and multidimensional scaling. The authors drew upon their extensive knowledge of the field and offered an insightful interpretation of 12 specialties identified in terms of 12 factors. The most well-known finding of the study is that information science at the time consisted of two essentially independent camps, namely, the information retrieval camp and the literature camp, including citation analysis, bibliometrics, and scientometrics.

Zhao and Strotmann (2008a, 2008b) followed up White and McCain's study using the same set of 12 journals and the same number of 120 cited authors in an updated time frame of 1996-2005. ... Zhao and Strotmann (2008b) found five major specialties and manually labeled them as user studies, citation analysis, experimental retrieval, Webometrics, and visualization of knowledge domains. In contrast to the findings of (White & McCain, 1998), experimental retrieval and citation analysis retained their fundamental roles in the field, and the user studies specialty became the largest specialty. Webometrics and visualization of knowledge domains appeared to make connections between the retrieval camp and the citation analysis camp.

A DCA by Aström (2007) studied papers published between 1990 and 2004 in 21 library and information science journals. Results were depicted in multidimensional scaling (MDS) maps. Aström's study also identified the two-camp structure found by (White & McCain, 1998). On the other hand, Aström found an information seeking and retrieval camp, instead of the information retrieval camp as in (White and McCain).

Although manually labeling a cocitation cluster can be a very rewarding process of learning about the underlying specialty and result in insightful and easy to understand labels, it requires a substantial level of domain knowledge and it tends to be time-consuming and cognitively demanding because of the synthetic work required over a diverse range of individual publications.

Traditionally, researchers often identify the nature of a cocitation cluster based on common themes among its members. ... The emphasis on common areas is a practical strategy; otherwise, comprehensively identifying the nature of a specialty can be too complex to handle manually.

Many researchers have studied the structural and dynamic properties of specialties in information science in terms of clusters, multivariate factors, and principle components (Morris & Van der Veer Martens, 2008; Persson, 1994; Tabah, 1999; White & Griffith, 1982).

A recent study of information science (Ibekwe-SanJuan, 2009) mapped the structure of information science at the term level using a text analysis system TermWatch and a network visualization system Pajek, but it did not address structural patterns of cited references.

Researchers also studied the structure of information science qualitatively, especially with direct inputs from domain experts. For example, Zins conducted a Critical Delphi study of information science, involving 57 leading information scientists from 16 countries (Zins, 2007a, 2007b, 2007c, 2007d).

Janssens, Leta, Glänzel, and De Moor (2006) studied the full-text of 938 publications in five library and information science journals with latent semantic analysis (LSA; Deerwester, Dumais, Landauer, Furnas, & Harshman, 1990) and agglomerative clustering. They found an optimal 6-cluster solution in terms of a local maximum of the mean silhouette coefficients (Rousseeuw, 1987) and a stability diagram (Ben-Hur, Elisseeff, & Guyon, 2002). Their clusters were labeled with single-word terms selected by tf*idf (p. 1625), which are not as informative as multiword terms for cluster labels.

Klavans, Persson, and Boyack (2009) recently raised the question of the true number of specialties in information science. They suspected that the number is much more than the 11 or 12 as reported in ACA studies such as (White & McCain, 1998) and (Zhao & Strotmann, 2008a, 2008b), but significantly fewer than the 72 reported in their own study, which is also based on the 12 journals between 2001 and 2005.

The 12-journal Information Science dataset, retrieved from the Web of Science, contains 10,853 unique bibliographic records, written by 8,408 unique authors from 6,553 institutions and 89 countries. These articles cited 129,060 unique references for a total of 206,180 times. They cited 58,711 unique authors and 58,796 unique sources.

The traditional procedure of cocitation analysis for both DCA and ACA comprises the following steps:
1. Retrieve citation data from sources such as the Science Citation Index (SCI), Social Science Citation Index (SSCI), Scopus, and Google Scholar.
2. Construct a matrix of cocited references (DCA) or authors (ACA).
3. Represent the cocitation matrix as a node-and-link graph or as a multidimensional scaling (MDS) configuration with possible link pruning using Pathfinder network scaling or minimum spanning tree algorithms.
4. Identify specialties in terms of cocitation clusters, multivariate factors, principle components, or dimensions of a latent semantic space using a variety of algorithms for clustering, community finding, factor analysis, principle component analysis, or latent semantic indexing.
5. Interpret the nature of cocitation clusters.

The interpretation step is the weakest link. It is time-consuming and cognitively demanding, requiring a substantial level of domain knowledge and synthesizing skills. In addition, much of attention routinely focuses on cocitation clusters per se, but the role of citing articles that are responsible for the formation of such cocitation clusters may not be always investigated as an integral part of a specialty.

Our new method extends and enhances traditional cocitation methods in two ways: (a) by integrating structural and content analysis components sequentially into the new procedure and (b) by facilitating analytic tasks and interpretation with automatic cluster labeling and summarization functions. The new procedure is highlighted in yellow in Figure 2, including clustering, automatic labeling, summarization, and latent semantic models of the citing space (Deerwester et al., 1990).

Our new procedure adopts several structural and temporal metrics of cocitation networks and subsequently generated clusters.

Structural metrics include betweenness centrality, modularity, and silhouette.

Temporal and hybrid metrics include citation burstness and novelty

The betweenness centrality metric is defined for each node in a network. It measure the extent to which the node is in the middle of a path that connects other nodes in the network (Brandes, 2001; Freeman, 1977). High betweenness centrality values identify potentially revolutionary scientific publications (Chen, 2005) as well as gatekeepers in social networks.

In the context of this study, the modularity Q measures the extent to which a network can be divided into independent blocks, i.e., modules (Newman, 2006; Shibata, Kajikawa, Taked, & Matsushima, 2008).

The silhouette metric (Rousseeuw, 1987) is useful in estimating the uncertainty involved in identifying the nature of a cluster.

Burst detection determines whether a given frequency function has statistically significant fluctuations during a short time interval within the overall time period.

Sigma is introduced in (Chen, et al., 2009a) as a measure of scientific novelty. ... In this study, Sigma is defined as (centrality + 1)burstness such that the brokerage mechanism plays more prominent role than the rate of recognition by peers.

We adopt a hard clustering approach such that a cocitation network is partitioned to a number of nonoverlapping clusters.

In this article, cocitation similarities between items i and j are measured in terms of cosine coefficients.

A good partition of a network would group strongly connected nodes together and assign loosely connected ones to different clusters. This idea can be formulated as an optimization problem in terms of a cut function defined over a partition of a network. Technical details are given in relevant literature (Luxburg, 2006; Ng, Jordan, & Weiss, 2002; Shi & Malik, 2000).

Spectral clustering is an efficient and generic clustering method (Luxburg, 2006; Ng et al., 2002; Shi & Malik, 2000). It has roots in spectral graph theory. Spectral clustering algorithms identify clusters based on eigenvectors of Laplacian matrices derived from the original network.

Spectral clustering has several desirable features compared to traditional algorithms such as k-means and single linkage (Luxburg, 2006):
 • It is more flexible and robust because it does not make any assumptions on the forms of the clusters,
• it makes use of standard linear algebra methods to solve clustering problems, and
• it is often more efficient than traditional clustering algorithms.

Candidates of cluster labels are selected from noun phrases and index terms of citing articles of each cluster. These term are ranked by three different algorithms. In particular, noun phrases are extracted from titles and abstracts of citing articles. The three term ranking algorithms are tf*idf (Salton, Yang, & Wong, 1975), log-likelihood ratio (LLR) tests (Dunning, 1993), and mutual information (MI).

Each cocitation cluster is summarized by a list of sentences selected from the abstracts of articles that cite at least one member of the cluster.

In this study, sentences are ranked by Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007). Given a set S of N sentences, let M be the square matrix that for each pair of sentences gives the number of nominal words in common (nouns and adjectives).

In this study, summarization sentences were also ranked by two new functions gtf and gftidf , which are further simplified approximations of the energy function E.

The ACA and DCA studies described in this article were conducted using the CiteSpace system (Chen, 2004; Chen, 2006). CiteSpace is a freely available Java application for visualizing and analyzing emerging trends and changes in scientific literature.

CiteSpace supports a unique type of cocitation network analysis—progressive network analysis—based on a time slicing strategy and then synthesizing a series of individual network snapshots defined on consecutive time slices. Progressive network analysis particularly focuses on nodes that play critical roles in the evolution of a network over time. Such critical nodes are candidates of intellectual turning points.

In summary, (a) spectral clustering and factor analysis identified about the same number of specialties, but they appeared to reveal different aspects of cocitation structures and (b) cluster labels chosen from citers of a cluster tend to be more specific terms than those chosen by human experts.

We found the comparison with the study of Zhao and Strotmann very valuable. It offered us an opportunity to compare the analysis conducted by human experts to the interpretation cues provided by our automatic labeling and summarization methods.

Spectral clustering for the purpose of network decomposition is exclusive in nature although in reality it is often sensible to allow overlapping clusters because of multiple roles individual entities may play.

Spectral clustering of cocitation networks tends to generate distinct clusters with high precision, whereas human experts tend to aggregate entities into broadly defined clusters.

In conclusion, the new cocitation analysis procedure has the following advantages over the traditional one:
• It can be consistently used for both DCA and ACA.
• It uses more flexible and efficient spectral clustering to identify cocitation clusters.
• It characterizes clusters with candidate labels selected by multiple ranking algorithms from the citers of these clusters and reveals the nature of a cluster in terms of how it has been cited.
• It provides metrics such as modularity and silhouette as quality indicators of clustering to aid interpretation tasks.
• It provides integrated and interactive visualizations for exploratory analysis.

Modularity and silhouette metrics provide useful quality indicators of clustering and network decomposition.

2015年3月24日 星期二

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

由於發現新的物種、疾病與社交模式,產生新的研究主題與專業 (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012),經過一段時間後,相關的研究社群會成長或是規模改變,有些主題仍然持續,但有些則是消失 (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010)。已有許多研究利用書目資料來確認研究的專業,例如Kessler (1963)的論文書目耦合網路(bibliographic coupling networks)、Small (1973)的論文共被引網路(paper co-citation networks)、White 與 McCain (1998)的作者共被引網路(author co-citation networks)以及White (2003)的尋路網路 (pathfinder networks),Callon、Courtial 與 Laville (1991)、Ding、Chowdhury 與 Foo (2000)、Milojevic、Sugimoto、Yan 與 Ding (2011)則是使用詞語共現網路 (co-word networks)。這些研究各自在不同研究層次確認研究主題,例如論文層次有 Chen (2004, 2006)、 Kessler (1963)和 Small (1973),作者層次有 Clauset, Newman, & Moore (2004)、White & McCain (1998)和 White (2003),期刊層次如 Glänzel & Schubert (2003)、 Leydesdorff & Vaughan (2006),以及領域層次有 Janssens, Zhang, Moor, & Glänzel (2009)、Rafols & Leydesdorff (2009)、Zhang, Liu, Janssens, Liang, & Glänzel (2010)。較低的研究實體層級,如論文與作者,研究可以從領域內發現其他的主題或專業;但在期刊或領域等較高的層次,通常從更完整的資料中確認出次領域。確認主題的方法則有因素分析(factor analysis)和多維尺度(multidimensional scaling)等傳統的群集技術以及連結線中心性(edge betweenness)、群組性(modularity)和混合群集(hybrid clustering)等較新技術的應用。本研究(Yan, 2014)則是利用主題模型(topic model)確認研究主題,並提出主題延續性(topic continuity)及主題普遍性(topic popularity)等兩項動態特性來分析研究主題。應用主題模型技術考察主題動態的方法,包括事後分析(post hoc analysis)(例如: Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008)、分段法(segmented approaches) (例如:Bolelli, Ertekin,Zhou, & Giles, 2009)以及連續時間模型(continuous-time model) (Wang & McCallum, 2006)等。本研究採用事後分析,利用文件中各主題的機率分布評估主題存在的機率。

本研究針對每一年分別產生一個主題模型,對於每一個主題找出後一年最有可能的主題,評估兩個主題相似的方式是利用改良自Kullback-Leibler差異 (Kullback-Leibler divergence, KLD)的Jensen–Shannon差異 (Jensen–Shannon divergence, JSD),兩個模型P和Q的JSD計算方式為JSD(P||Q) = 1/2KLD(P||M)+1/2KLD(Q||M),KLD是兩個模型Kullback-Leibler差異,M=1/2(P+Q)。每一個主題後一年最有可能的主題是擁有最小JSD的主題,其分數JSD稱為JJSDS,運用JJSDS的變化趨勢計算主題的連續性,然後以z score進行標準化。
評估各主題的普遍性則是計算它們在該年度文件上平均的機率值,並以z score進行標準化,愈大的機率值表示該主題在當年度愈普遍,然後分析主題普遍性的變化趨勢,。

本論文的研究資料為2001到2011年的圖書資訊學(library and information science)出版品,包括期刊論文、書評及研討會論文等,採用論文的題名做為分析資料,共27,796 篇論文。每一年的主題數目都設為20。結果顯示在網路資訊檢索(web information retrieval)、引用及書目計量學(citation and bibliometrics)、系統及技術(system and technology)、健康科學(health science)等主題有較高的平均普遍性;h指標(h-index)、線上社群(online communities)、資料保存(data preservation)、社群媒體(social media)和網站分析(web analysis)等則是圖書資訊學裡愈來愈普遍的主題。研究結果的主題與過去的研究相符合,但這篇論文的貢獻在於對於研究主題的動態進行分析。

Dynamic development is an intrinsic characteristic of research topics. To study this, this paper proposes two sets of topic attributes to examine topic dynamic characteristics: topic continuity and topic popularity.

Topic continuity comprises six attributes: steady, concentrating, diluting, sporadic, transforming, and emerging topics; topic popularity comprises three attributes: rising, declining, and fluctuating topics.

These attributes are applied to a data set on library and information science publications during the past 11 years (2001–2011).

Results show that topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity; topics on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular in library and information science.

Dynamics is a constant theme in scientific explorations. Research communities may grow or change in size; new species, diseases, or societal patterns may be discovered; and new research topics and specialties may be introduced (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012). Over time, some topics are continuously investigated while others appear or disappear (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010). Therefore, it is of great importance to examine research dynamics to understand the evolving cognitive structures of research domains.

Pioneering studies of paper bibliographic coupling networks (Kessler, 1963), paper co-citation networks (Small, 1973), author co-citation networks (White & McCain, 1998), pathfinder networks (White, 2003) and co-word networks (e.g., Callon, Courtial, & Laville, 1991; Ding, Chowdhury, & Foo, 2000; Milojevic, ´ Sugimoto, Yan, & Ding, 2011) were capable of identifying research specialties from bibliographic data effectively.

However, findings from these studies remained largely static and thus only yielded fixed perspectives on the cognitive structure of research domains.

To examine research dynamics,this study uses a topic modeling technique and proposes two sets of topic attributes–topic continuity and topic popularity.

• How to use topic modeling techniques to study research dynamics?
• What quantitative measurements can be used to describe topic dynamics?
• What topics are present in library and information science? What are their dynamic characteristics?

This subsection reviews the network-based approaches of identifying research topics and specialties. These approaches have been applied to several research levels, including the paper-level (e.g., Chen, 2004, 2006; Kessler, 1963; Small, 1973), the author-level (e.g., Clauset, Newman, & Moore, 2004; White & McCain, 1998; White, 2003), the journal-level (e.g., Glänzel & Schubert, 2003; Leydesdorff & Vaughan, 2006), and the field-level (e.g., Janssens, Zhang, Moor, & Glänzel, 2009; Rafols & Leydesdorff, 2009; Zhang, Liu, Janssens, Liang, & Glänzel, 2010).

Most above-mentioned work used co-occurrence networks as the research instrument.

Analyses on lower level research entities, such as papers and authors, usually identified topics and specialties from small but well-defined research fields; whereas analyses on higher level research entities, such as journals and fields, attempted to identify subfields and subdomains from more comprehensive data sets.

Both classic clustering techniques (e.g., factor analysis and multidimensional scaling) as well as modern techniques (e.g., edge betweenness, modularity, and hybrid clustering) have been applied.

Recently, studies have attempted to add dynamic analyses by utilizing multiple time intervals.

Several approaches on slicing time intervals are available: intervals that have the same amount of references (e.g., Radicchi et al., 2009), intervals that have the same number of publications (e.g., Sugimoto, Li, Russell, Finlay, & Ding, 2011; Yan & Sugimoto, 2011), same-length intervals (e.g., Åström, 2007; Milojevic´ et al., 2011), and accumulative intervals (e.g., Barabási et al., 2002; Yan & Ding, 2009).

These studies laid valuable methodological basis for dynamic analyses of cognitive structures of research fields; however, networks of different time frames were largely analyzed distinctively and a more integrated examination was lacking.

In the meantime, empirically, network-based clustering results may require domain expertise to effectively interpret obtained results.

Topic modeling techniques use probabilistic models to assign papers, journals, or authors to clusters. A topic can be defined as a probability distribution over terms in a vocabulary (Blei & Lafferty, 2007). Latent Dirichlet Allocation (LDA) model, a classic topic model, was proposed by Blei et al. (2003). The model predicates that words for each paper are derived from a mixture of topics and each topic follows a multinomial distribution.

One recent update of the LDA model is the supervised LDA model. It makes the analyses of multi-labeled corpora (e.g., tags from delicious.com and various classifications) possible. Blei and McAuliffe’s (2010) version of supervised LDA can successfully address this challenge, but a document can only be assigned with one label.

Ramage, Hall, Nallapati, and Manning (2009) offered an approach which enabled the multi-label assignment. Their supervised labeled LDA (L-LDA) associated one label with one topic and allowed the model to learn word-label relations.

Through topic modeling techniques, topic dynamics has been examined mainly through the following approaches: post hoc analysis (e.g., Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008), segmented approaches (e.g., Bolelli, Ertekin,Zhou, & Giles, 2009), and continuous-time model (Wang & McCallum, 2006).

Post hoc analysis uses topic-document probability distributions to evaluate the presence of identified topics.

Segmented approaches build the dynamic component in the probabilistic model. It assumes that the state of topics at a single time point is independent from all other time points and divides document corpora into segments that have contingent time stamps (Bolelli et al., 2009).

Continuous-time model is a non-Markov model proposed by Wang and McCallum (2006), where they found the non-Markov model provides better prediction and more interpretable topical trends.

In this study, a post hoc dynamic analysis using the ACT model is selected because of its marked performance (Tang et al., 2008) as well as its advanced input and output support.

Topic dynamics is calculated through the Author-Conference-Topic (ACT) model (Tang et al., 2008).

Specifically, i is the topic distribution for document i. Mean ( ¯), therefore, is a direct quantitative measurement to assess topic popularity: the higher the ¯, the more visible the topic, and thus the more popular that topic is (Griffiths & Steyvers, 2004).

Because the data set spans 11 years, 11 independent ACT models were run, one for each year of the data set based on year of publication.

The Jensen–Shannon divergence (JSD) was used as the similarity measurement to quantify the topic similarity between different word-topic distributions. ... JSD is a symmetrized and smoothed version of the Kullback–Leibler divergence (KLD). ... As a divergence measure, the smaller the JSD, the higher the similarity is.

In order to track the same topic from two adjacent time intervals, the minimum value for each row of a JSD matrix was used, referred to as the joint JSD score (JJSDS): MIN(JSD Matrix(i,j)), for j = 1:n. ... Applying the same approach to each pair of adjacent time slices, for each topic, an array of JJSDS can be obtained.

The attributes of steady, concentrating, and diluting topics focus on the overall topical characteristics whereas the attributes of sporadic, transforming, and emerging topics focus on the topical characteristics of a specified time frame. Therefore, these attributes are not mutual exclusive, suggesting that a topic can be a concentrating topic overall, and in the meantime, related topics were added and thus qualifying it for a transforming topic.

The data set contains publications of all journals indexed in the 2011 version of the Journal Citation Report in the Information Science & Library Science subject category. Articles, proceeding papers, and review articles published within these journals from 2001 to 2011 were downloaded for analysis (downloading time: October 2012). Stop words were then removed from publications’ titles. Publications without titles, authors, or journal names were removed from the data set. The final data set comprised 27,796 papers.

The number of topics is set at 20: this number considers the size of the paper corpus as well as previous empirical studies on the cognitive structure of library and information science (e.g., Milojevic´ et al., 2011; Sugimoto et al., 2011; White & McCain, 1998; Zhao & Strotmann, 2008). For reasons of consistency, the same number of topics was identified for each year of the data set.

In this subsection, we first present histograms made from values in Jensen–Shannon divergence (JSD) matrices (Fig. 4). These histograms provide a direct perception on how research topics in library and information science are related as measured by JSD. This subsection then introduces all 20 topics in each year from 2001 to 2011 as well as how topic continuity and popularity attributes are applied to these topics (Fig. 5).

Fig. 4 uses histograms to visualize JSD values for each pair of adjacent years. Because there are 20 topics for each year, the number of data points in each histogram is 400 (20 × 20). This number is 4000 for the histogram in the lower right section of Fig. 4, as it uses JSD values for all pairs of adjacent years.

This study finds that in library and information science, research topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity over the past decade (from 2001 to 2011).

Research on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular topics.

Overall, findings of this study are consistent with previous studies using co-word, co-citation, and topic modeling techniques.

For instance, a co-word study by Milojevic´ and colleagues (2011) has found that title terms “citation”, “impact factor”, and “web” have a rising usage from 1989 to 2008.

Other related dynamic studies that cover the target time frame of the current study (2001–2011) include Åström’s (2007) study on examining library and information science research front, where the study found that webometrics and information-seeking and retrieval have become dominating research areas between 2000 and 2004.

This finding has been verified by Klavans and Boyack (2011) where the authors used the global map (i.e., the map of science) to enhance to accuracy of local maps (i.e., the contextual map of information science). They identified five core areas in information science, including information-seeking behavior, computer-enhanced retrieval, scientometrics, co-citation analysis, and citation behavior.

Besides the contextual analysis of information science, structural analysis has also been achieved from a time-series empowered author co-citation and document co-citation analysis (Chen, Ibekwe-SanJuan, & Hou, 2010). Through the application of a series of structural metrics such as centrality measures, modularity and silhouette, a clear cognitive structure of information science was attained in that the research areas of interactive information retrieval, academic web, information retrieval, citation behavior, and h-index have gained a particular popularity from 1996 to 2008.

In addition to journal publications, Sugimoto and colleagues (2011) applied a LDA model to library and information science dissertations and demonstrated dissertations as an important communicative genre. Their study indicated that between 2000 and 2009, internet and information retrieval related topics were the central dissertation research themes.

The contribution of the current study is that it proposes two sets of quantitative topic attributes. These attributes have streamlined the dynamic analysis of research topics and specialties and have further complemented co-occurrence-based studies.

This paper has identified dynamic characteristics of topics in library and information science; however, limited information can be told about the mechanisms that resulted in such characteristics. That being said, the study is unable to pinpoint, for instance, whether the growing popularity of network and citation studies is the result of a growing research community, a drive by the commercial market, a stimulus from funding agencies, or a combination of these or other unlisted factors.

Popular topics may be associated with research communities that are expanding in size and/or tend to have higher productivity. Conversely, less popular topics may be associated with communities that are shrinking and/or have a reduced productivity. Topic continuity and popularity attributes reflect research specialties’ development in scientific communities, which is further guided by science policies and the attention of the general public.

In informetrics, studies have mainly focused on analyzing the performance and the social and cognitive implications of several types of research entities, including papers, authors, institutions, journals, and fields. Authors and institutions are typically used to examine social relations in academia; while journals and fields are predominantly used to investigate the cognitive structure of research domains.

Topic analysis can precisely provide a more refined assessment by clustering research papers based on certain probability distributions. Because of such quantitative results, a more integrated dynamic cognitive analysis is thus possible, as exemplified through the current study.

Topic analysis will be further developed by overlaying topics with author communities to explore the interwoven relationships between research topics and research communities (e.g., Yan et al., 2012); by overlaying topics with funding data to investigate the “lead-lag” relationship between funding support and productivity (e.g., Shi et al., 2010); by applying topic models to different genres to study research immediacy (e.g., Ding et al., 2013); and by overlaying topics with citation data to examine the relationships between topics and impact.

2014年9月15日 星期一

Bonnevie-Nebelong, E. (2006). Methods for journal evaluation: journal citation identity, journal citation image and internationalization. Scientometrics, 66(2), 411-424.

Bonnevie-Nebelong, E. (2006). Methods for journal evaluation: journal citation identity, journal citation image and internationalization. Scientometrics, 66(2), 411-424.

Scientometrics

本研究以引用分析對Journal of Documentation (J DOC) 進行評估,並且與JASIST和JIS進行比較。所使用的引用分析方法包括三個方面:以引用的參考文獻為主的期刊引用認同(journal citation identity)、以被引用的情形為主的期刊引證形象(journal citation image)和以出版品本身為主的國際化(internationalisation)。

在期刊引用認同方面有兩種指標。第一種指標是引用對被引用者比(citations/citee-ratio),計算方式是分析範圍內所有參考文獻數除以參考文獻上出現的期刊種類,如果這個數值愈低,表示出現許多不同種類的期刊,也就是使用的期刊具有多元性(diversity)。另一個指標是自我引用(self-citations),用來測量期刊在科學領域內的獨立性(isolation),如果自我引用的程度低表示在科學領域內的影響力高。自我引用指標的測量包括引用文獻中來自本身期刊的比例(self-citing)和期刊被引用的情形下來自本身的比例(self-cited),前者是期刊引用認同的一部份,而後者則屬於期刊引證形象。

期刊引證形象也包含兩種指標。第一種指標是新期刊擴散因素(new journal diffusion factor),此一指標分析該期刊時間範圍內每一篇論文平均被引用的期刊種類,代表該期刊的想法出口情形(export of ideas)、跨領域性(transdisciplinarity)以及專殊化(specialisation)程度。另一只標示該期刊的共被引期刊,根據共被引情形以及共被引期刊的期刊影響因素(Journal Impact factor)來加以描述。

國際化是測量出版品以及引用期刊論文的作者地區。

各種分析方法與指標整理為Table 1。



首先,引用對被引用者比的結果如Figure 1。另外,1990到2003年的平均引用對被引用者比,JDOC為1.50,JASIST為1.88,JIS則為1.44。較低的引用對被引用者比表示引用的文獻裡重複的期刊較多種,代表這份期刊有較多元的科學基礎(scientific base)。從結果上看來,JDOC比JASIST的科學基礎多元性較高,但較JIS來得低。

JDOC比JASIST和JIS的文章有較高的比例是書評(boo review),這使得JDOC的參考文獻數較少,因為書評平均只有1.6到2筆參考文獻。

在1990到2003年間,JDOC、JASIST和JIS等三種期刊引用本身的比例都有下降的趨勢,表示這三種期刊愈來愈不孤立,測量期刊被引用的情形,則可發現JDOC與JIS來自期刊本身的比例則較低,表示它們在這個領域的能見度(visibility)較高。另外,JIS在從1979年開始的前十年引用來自期刊本身的比例較高,則說明了這個期刊在當時為在領域邊緣的新期刊。

JDOC的新期刊擴散因素比其他兩種期刊稍大,並且有往上的趨勢。

經常與JDOC共同引用的前十種期刊如Table 3所示。期刊共被引的相似度以Jaccard Similarity測量。

JDOC上論文作者的地區分布如Figure 10。主要的作者來自西歐地區,並逐漸增加。



引用JDOC論文的作者地區分布則如Figure 11。以北美地區的作者引用最多,但西歐地區則逐漸增加。



The Journal Citation Identity is a reference analysis. It is measured by looking into  the referencing style of the publishing authors. What is their combined citations/(journal) citee-ratio? This means that the total number of references in the journal must be calculated, year-by-year or all years taken together. The result of this is divided with the number of different journals present in the set of references. If the set contains many different journals, the ratio will be lower. Consequently a low average signifies a greater diversity in the use of journals among the authors as part of their scientific base, and thus a wider horizon.

Self-citations are part of the Journal Citation Identity as well as the Journal Citation Image, depending on the perspective. ... They are indicators of the style of a journal. Many self-citations among the references may signify isolation of the journal in the scientific domain (high rate). A low rate of self-citations may indicate a high level of influence in the scientific community.

The Journal Citation Image is based on citation analyses of two types: the New Journal Diffusion Factor (N JDF) and journal co-citation analysis.

The New Journal Diffusion Factor was proposed by Frandsen, and is inspired by Rowlands’ diffusion factor. It measures breadth by number of citing journals per published article. N JDF is the average number of different journals that an average article is cited by within a given time window. The result of this tells about the scientific style and about breadth, export of ideas, transdisciplinarity and degree of specialisation of a journal. N JDF is tested for JDOC in a time perspective.

The Journal Citation Image “the White way” means to do a co-citation journal-by-journal analysis and interpret the result in a qualitative manner. It is thus a means to evaluate a journal by the journals co-cited with the journal in question. The co-cited journals are displayed in a list ranked by frequency of co-incidences, the number of citations for each co-cited journal taken into consideration by application of the jaccard calculations. Also the Journal Impact factor (JIF) is used to evaluate the co-cited journals. The co-cited journals then function as image-makers of the journal in question.

Internationalisation is measured by looking into the geographic locations of both publishing and cited authors of the JDOC.

A high citation/citee ratio means that the journal has many recited journals among its references. A low ratio signifies less journal re-citations and thus a greater diversity of journals as part of the scientific base and a wider horizon among authors.

Journal self-citations. Journal self-citations can be analysed from two perspectives, by self-citing rate and by self-cited rate. The first mentioned is part of the citation identity, the second one is part of the self-image, but the two types of self-citations are treated together here for practical reason.

The three journals all show decreasing self-citing rates during the years 1980–2003. This may signify a tendency towards less isolation of the field.