顯示具有 factor analysis 標籤的文章。 顯示所有文章
顯示具有 factor analysis 標籤的文章。 顯示所有文章

2015年4月15日 星期三

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V. (2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V.(2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

由於一般認為將領域之間的關係表示為圖形,通過考慮這些關係的可能性能夠提供許多資訊,不論對新進人員或專家皆有助於理解與分析,因此對這方面方法與工具的需求逐漸提高。過去的研究大多以期刊為分析單位,產生所有科學研究領域的科學映射圖。例如Leydesdorff (2004a, 2004b)使用雙重連結成分(biconnected components)的圖形分析演算法,將JCR 2001的科學研究進行分類。Boyack, Klavans, and Börner (2005)則應用了8種不同的期刊相似性測量7121種SCI和SSCI期刊,並採用VxOrd產生科學映射圖。Samoylenko, Chao, Liu, and Chen (2006)建構科學期刊的最小生成樹(minimum spanning trees),他們使用的資料是SCI 1994到2001的資料。本研究提出一個將ISI (Institute of Scientific Information)類別繪製成科學映射圖的方法,這個方法利用根據類別間的共被引資訊建構類別間的連結,以尋徑網路(PathfinderNetwork)縮減不重要的連結,然後以Kamada-Kawai方法決定節點在圖上的布局(layout),最後利用因素分析(factor analysis)進行結構確認。本研究和先前的研究都是針對類別利用共被引資訊呈現科學映射圖。以類別為分析單位在代表上足夠明確,並且比起較小的單位,這種方式對非專家使用者(nonexpert user)較具有資訊且使用者友善。Moya-Anegón et al. (2004)針對西班牙科學研究領域的視覺化,Moya-Anegón et al. (2005)則進一步利用科學映射圖比較英國、法國和西班牙三個國家的科學研究領域。本研究依循Börner, Chen, and Boyack (2003)提出的知識領域映射流程。使用的資料為7585種ISI期刊,ISI的類別共有219個,但扣除多學科科學後(Multidisciplinary Sciences),採用的類別共218個。利用共被引計算期刊相似性的方式為

Cc(ij)為期刊i和期刊j共被引次數,c(i)和c(j)則分別是期刊i和期刊j被引用次數。然後以尋徑網路和Kamada-Kawai方法繪製網路圖,經過尋徑網路處理後,有較多連結的節點具有較重要的地位。而尋徑網路是一種以型態為主的方法,與以群集為主的因素分析彼此間可以互補,因素分析可以識別、界定與定名科學映射圖上呈現的主題區域,而尋徑網路則負責讓使主題區域更加明顯,將類別分組成束,並顯示連接不同顯著類別的路徑,以及總體的型態結構。。最後總計共分析出35個因素,通過陡坡考驗(scree test)則有16個。科學映射圖上的類別可以分為三個群集:醫學與地球科學、基礎與實驗科學以及社會科學。

This study proposes a new methodology that allows for the generation of scientograms of major scientific domains, constructed on the basis of cocitation of Institute of Scientific Information categories, and pruned using PathfinderNetwork, with a layout determined by algorithms of the spring-embedder type (Kamada–Kawai), then corroborated structurally by factor analysis.

We present the complete scientogram of the world for the Year 2002.

This need arises from the general conviction that an image or graphic representation of a domain favors and facilitates its comprehension and analysis, regardless of who is on the receiving end of the depiction and whether a newcomer or an expert.

Science maps can be very useful for navigating around in scientific literature and for the representation of its spatial relations (Garfield, 1986). They are optimal means of representing the spatial distribution of the areas of research while also offering additional information through the possibility of contemplating these relationships (Small & Garfield, 1985).

From a general viewpoint, science maps reflect the relationships between and among disciplines; but the positioning of their tags clues us into semantic connections while also serving as an index to comprehend why certain nodes or fields are connected with others.

Moreover, these large-scale maps of science show which special fields are most productively involved in research—providing a glimpse of changes in the panorama—and which particular individuals, publications, institutions, regions, or countries are the most prominent ones (Garfield, 1994).

It is a tool in that it allows the generation of maps, and a method in that it facilitates the analysis of domains, by showing the structure and relations of the inherent elements represented. In a nutshell, scientography is a holistic tool for expressing the discourse of the scientific community it aspires to represent, reflecting the intellectual consensus of researchers on the basis of their own citations of scientific literature.

In Moya-Anegón et al. (2004), we ventured forth with a historic evolution of scientific maps from their origin to the present, and proposed ISI-JCR category cocitation for the representation of major scientific domains. Its utility was demonstrated by a visualization of the scientific domain of geographical Spain for the Year 2000.

Since then, other works related with the visualization of great scientific domains have appeared; however, all use journals as the unit of analysis, with the exception of a study based on the cocitation of categories (Moya-Anegón et al., 2005), comparatively focusing on three geographic domains (England, France, and Spain).

In contrast, Leydesdorff (2004a, 2004b) classified world science using the graph-analytical algorithm of biconnected components in combination with JCR 2001.

Boyack, Klavans, and Börner (2005) applied eight alternative measures of journal similarity to a dataset of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes, to show the first global map of science using the force-directed graph layout tool VxOrd.

Samoylenko Chao, Liu, and Chen (2006) proposed an approach through the construction of minimum spanning trees of scientific journals, using the Science Citation Index from 1994 to 2001.

In processing and depicting the scientific structure of great domains, we further developed a methodology that follows the flow of knowledge domains and their mapping as proposed by Börner, Chen, and Boyack (2003).

Because ISI assigns each journal to one or more subject categories, to designate a subject matter (i.e., ISI category) for each document, we also downloaded the Journal Citation Report (JCR; Thomson Corporation, 2005a), in both its Science and Social Sciences editions, for 2002.

The downloaded records were exported to a relational database that reflects the structured information of the documents. This new repository contained nearly 1 million (N = 901,493) source documents: articles, biographical items, book reviews, corrections, editorial materials, letters, meeting abstracts, news items, and reviews that had been published in 7,585 ISI journals (N = 5,876 + 1,709). These were classified in a total of 219 categories, altogether citing 25,682,754 published documents.

As informational units, they are, in themselves, sufficiently explicit to be used in the representation of all disciplines that make up science in general. These categories, in combination with the adequate techniques for the reduction of space and the representation of the information to construct scientograms of science or of major scientific domains, prove much more informative and user friendly for quick comprehension and handling by nonexpert users than those obtained by the cocitation of smaller units of cocitation.

For these reasons, we used the 219 categories of the JCR 2002 as units of measure, with the exception of “Multidisciplinary Sciences.” ... The maximum number of categories with which we worked, then, was 218.

In light of our previous experience (Moya-Anegón et al., 2004, 2005), we use cocitation as the similarity measure to quantify the relationship existing between each one of the JCR categories.

Therefore, after a number of trials, we arrived at the conclusion that using tools of Network Analysis, the best visualizations are those obtained through raw data cocitation as the unit of measure. Yet, it also was necessary to reduce the number of coincident cocitations to enhance pruning algorithm yield. Therefore, to those raw data values we added the standardized cocitation value. In this way, we could work with raw data cocitation while also differentiating the similarity values between categories with equal cocitation frequencies. The key was a simple modification of the equation for the standardization of the degree of citation proposed by Salton and Bergmark:




where CM is cocitation measure, Cc is cocitation frequency, c is citation, and i and j are categories.

Over the history of the visualization of scientific information, very different techniques have been used to reduce n-dimensional space. Either alone or in conjunction with others, the most common are multidimensional scaling, clustering, factor analysis, self-organizing maps, and PathfinderNetworks (PFNET).

In our opinion, PFNET with pruning parameters r = ∞, and q = n − 1 is the prime option for eliminating less significant relationships while preserving and highlighting the most essential ones, and capturing the underlying intellectual structure in a economical way.

Although PFNET has been used in the fields of Bibliometrics, Informetrics, and Scientometrics since 1990 (Fowler & Dearhold, 1990), its introduction in citation was due to the hand of Chen (1998, 1999), who introduced a new form of organizing, visualizing, and accessing information. The end effect is the pruning of all paths except those with the single highest (or tied highest) cocitation counts between categories (White, 2001).

The spring embedder type is most widely used in the area of documentation, and specifically in domain visualization. Spring embedders begin by assigning coordinates to the nodes in such a way that the final graph will be pleasing to the eye (Eades, 1984). Two major extensions to the algorithm proposed by Eades (1984) have been developed by Kamada and Kawai (1989) and Fruchterman and Reingold (1991).

While Brandenburg, Himsolt, and Rohrer (1995) did not detect any single predominating algorithm, most of the scientific community goes with the Kamada–Kawai algorithm. The reasons upheld are its behavior in the case of local minima, its capacity to minimize differences with respect to theoretical distances in the entire graph, good computation times, and the fact that it subsumes multidimensional scaling when the technique of Kruskal and Wish (1978) is applied.

We can effortlessly see which are the most important nodes in terms of the number of their connections and, in turn, which points act as intermediaries with other lines, as hubs or forking points.

Whereas factor analysis is a clustering-oriented procedure, PFNET is topology oriented. Yet, they are extremely valuable as complements in the detection of the structure of a scientific domain.

Thus, factor analysis is responsible for identifying, delimiting, and denominating the great thematic areas reflected in the scientogram.

Meanwhile, PFNET is in charge of making the subject areas more visible, grouping their categories into bunches, and showing the paths that connect the different prominent categories, and finally, the overall topology of the domain.

Factor analysis identifies 35 factors in the cocitation matrix of 218 × 218 categories of world science 2002. Through the scree test we extracted 16, which we tagged using the previously explained method; these accumulate 70.2% of the variance (Table 1)

The number of categories included in at least one factor is 195. Twenty-three were not included in any factor (Table 2), and 25 belonged to two factors simultaneously (Table 5).

That is, a category or thematic area occupying a central position in the scientogram will have a more general or universal nature in the domain as a consequence of the number of sources it shares with the rest, contributing more to scientific development than those with a less central position.

The more peripheral the situation of a category or subject area, the more exclusive its nature, and the fewer the sources it will appear to share with other categories; accordingly, the lesser its contribution to the development of knowledge through scientific publications.

An intermediary position favors the interconnection of other categories or thematic areas. 

This broad interpretation of our scientograms not only explains the patterns of cocitation that characterize a domain but also foments an intuitive way for specialists and nonexperts to arrive at a practical explanation of the workings of PFNET (Chen & Carr, 1999).

From a macrostructural point of view, we can distinguish three major zones.

In the center is what we could call Medical and Earth Sciences, consisting of Biomedicine, Psychology, Etiology, Animal Biology & Ecology, Health Care & Service, Orthopedics, Earth & Space Science, and Agriculture & Soil Sciences.

To the right, we can see some other basic and experimental sciences: Materials Sciences & Physics, Applied; Engineering; Computer Science & Telecommunications; Nuclear Physics & Particles & Fields; and Chemistry.

To the left is the neighborhood of the social sciences, with Applied Mathematics, Business, Law, and Economy, and Humanities.

On one hand, it offers domain analysts the possibility of seeing the most essential connections between categories of given domain.

On the other hand, it allows us to see how these categories are grouped in major thematic areas, and how they are interrelated in a logical order of explicit sequences.

2015年4月2日 星期四

Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

揭露與更好地了解科學傳播 (scholarly communication) 的中研究人員、研究團隊、機構、地區/國家、學科、出版品之間的關係,可以在不同尺度下進行研究。分析資料之間的連結可以從直接引用、共被引、共同著作、詞語或主題的共現,或者隱含式主題等形式進行。期刊間的相似性經常利用共被引,過去曾經進行期刊共被引研究的學科有經濟學 (McCain, 1991)、資訊檢索 (Ding, Chowdhury & Foo, 2000)、資訊系統 (Marion, Wilson & Davis, 2005)、醫療資訊學 (medical informatics, Morris & McCain, 1998)、類神經網路 (neural networks, McCain, 1998)以及半導體研究 (Tsay, Xu & Wu, 2003)。然而共被引研究有以下的困難:首先,在多個學科裡的應用,期刊與期刊間的共被引矩陣可能相當稀疏(Boyack, Klavans, & Börner, 2005)。其次,若是沒有Web of Science等資料庫來源,共被引資料的取得將會相當困難。最後,共被引分析大多只利用共被引次數,而沒有考慮引用來源的任何特性。

資訊計量學的重要研究之一是利用引用資料來確認期刊的學科或專業背景,錯誤的分類結果將會影響期刊在領域內的排名。Glänzel and Schubert (2003)發展一個三個步驟的期刊分類程序,使得期刊裡的文章可以根據參考文獻指定主題。Rafols and Leydesdorff (2009)對Web of Science的主題分類(Subject Categories)和Glänzel and Schubert (2003)的主題分類,比較兩種大型矩陣的分解演算法。Leydesdorff and Rafols (2009) 也使用主題分類引用頻率的引用矩陣研究170多種Web of Science的主題分類之間的關係。Leydesdorff and Schank (2008)以視覺化及動畫的方式呈現期刊之間的關係與它們的跨領域性。

過去的研究有利用期刊引用形象(journal citation image),也就是引用目標期刊的所有期刊的列表,做為期刊之間相似度評估的特徵。在計算上,可將期刊引用形象加入引用的頻率分布做為目標期刊的一種特徵(signature)。然而由於具有影響力與聲譽的期刊可能有相當大量的引用期刊,造成期刊引用形象的計算量較大。Wang and Wolfram (2014)提出使用引用期刊所屬的學科,利用引用學科的引用頻率做為期刊的特徵,來降低計算量。Wang and Wolfram (2014),提出引用學科分析(citing discipline analysis)來評估被引用期刊(cited journals)之間的相似性。引用學科分析根據目標期刊的引用期刊(citing journals)在Web of Science的研究領域上的頻率分布,取代直接。Wang and Wolfram (2014)並將引用學科分析應用於JCR的資訊科學與圖書館學 (Information Science & Library Science)的40種期刊,他們發現在多元尺度與群集分的結果中,若干期刊與其他期刊並不接近,有些同時被歸類於其他學科的期刊並不接近於資訊科學與圖書館學的期刊,而且當這些期刊同時也歸類於較大的相關領域時,其具有較高的影響係數(impact factors)將會降低其他期刊的排名。由於Wang and Wolfram (2014)只探討一個學科內的期刊,無法了解多個學科的期刊是否也具有同樣的情形。

本研究同樣利用引用學科分析估計期刊間的相似性。本研究使用來自於6個學科的120種期刊做為研究資料,其中5個學科彼此間較為接近,包括傳播學 (Communication)、電腦科學-資訊系統 (Computer Science-Information Systems)、教育學與教育研究 (Education & Educational Research)、資訊科學與圖書館學 (Information Science & Library Science)、管理學 (Management),另一個學科-地理學 (Geology) 則較遠。選取期刊的出版期間為1987到2012年,分為三個時期1987–1995、 1996–2004和 2005–2012。利用餘弦測量計算期刊間的相似性,並將估計結果應用於多元尺度(multidimensional scaling)、階層式群集分析(hierarchical cluster analysis)、主成分分析(Principal Component Analysis)等技術。

第一時期可以發現六個學科的相關期刊分為五個群集,其中電腦科學-資訊系統的期刊因為在這時期的數量較少而與資訊科學與圖書館學形成一個群集,地理學的群集與其他群集的距離相當遠。原本主題被分類在傳播學的 Journal of Advertising Research(JAR),在這時期的結果裡與管理學的相關期刊較接近。MIS相關期刊以及 Telecommunications Policy (TP)在資訊科學與圖書館學與管理學之間,Social Science Computer Review (SSCR)則是在資訊科學與圖書館學、傳播學與管理學三個學科間,另外Science Communication (SCOMM)雖然主題被分類在傳播學,但在本研究裡的結果,則是歸入資訊科學與圖書館學的群集。


第二時期電腦科學-資訊系統已經和與資訊科學與圖書館學分開,120種期刊共形成6個群集。在這時期的結果,部分主題分類於資訊科學與圖書館學的期刊被歸入其他學科,包括The Journal of the American Medical Informatics Association (JAMIA)和The International Journal of Geographical Information Science (IJGIS)被歸入電腦科學-資訊系統,後者甚至也很接近地理學;Decision Support Systems (DSS)則是原本在電腦科學-資訊系統分類下,而被歸入資訊科學與圖書館學的群集。Journal of Health Communication (JHC)的主題分類有傳播學和資訊科學與圖書館學,但在這時期更靠近傳播學的相關期刊而被歸類在傳播學的群集內。此外,原本分別在教育學與教育研究以及傳播學的the Academy of Management Learning& Education (AMLE)和JAR都被歸類在管理學的群集裡。

第三時期的六個群集對應到六個學科,但原先主題分類為傳播學的一些期刊更靠近於管理學期刊,而同時具有教育學與教育研究以及資訊科學與圖書館學兩種主題的International Journal of Computer-supported Collaborative Learning (IJCSCL),在結果上明顯可歸類為教育學與教育研究的群集。

有些被指定在某一個學科的期刊結果更接近其他學科,但有些被指定在多個學科的期刊則發現只有接近其中的一個學科。因此這個研究所提出的方法在計算期刊間的接近程度時能夠與期刊共被引分析(journal co-citation analysis)等傳統方法互補。由於以下的幾種原因,有愈來愈多的期刊跨越學科的邊界:1) 期刊出版的範圍愈來愈跨學科(more interdisciplinary),因此吸引愈來愈多其他學科的出版品引用。2) 整合式搜尋工具愈來愈普遍,更容易讓其他學科的作者發現。3) 愈多的期刊加入分析,使得期刊在引用學科的特徵上獨特性減少,彼此間更加相似。

為了更加了解跨學科的期刊,本研究利用主成分分析探討期刊的引用學科特徵。在第一時期,資訊科學與圖書館學的相關期刊中,共有五種期刊屬於兩種成分,包含ISJ, ISR, MISQ等三種管理學期刊以及屬於傳播學和電腦科學-資訊系統的SCOMM和DSS。但在資訊科學與圖書館學主題分類下的ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS以及 MISQ同時也被分類在電腦科學-資訊系統主題內,但卻沒有出現在電腦科學-資訊系統的成份裡。並且JAR和 AMLE分別被分類在傳播學和教育學與教育研究,但在本研究的第一時期結果只屬於管理學。

在第二時期的結果中,IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ等許多MIS期刊同時出現在管理學和資訊科學與圖書館學的成份裡。雖然SSCR僅有被分類在資訊科學與圖書館學主題,但同時包含在資訊科學與圖書館學與傳播學的成分中。另外,雖然IJGIS被分類在資訊科學與圖書館學主題,但並沒有出現在任何成分,顯然是六個學科較周邊的期刊。

第三時期同時在管理學和資訊科學與圖書館學成份裡的MIS期刊更增加了。沒有出現在任何成分的期刊,除了IJGIS以外,還增加了Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM)。

A similarity comparison is made between 120 journals from five allied Web of Science disciplines (Communication, Computer Science-Information Systems, Education & Educational Research, Information Science & Library Science, Management) and a more distant discipline (Geology) across three time periods using a novel method called citing discipline analysis that relies on the frequency distribution of Web of Science Research Areas for citing articles.

Similarities among journals are evaluated using multidimensional scaling with hierarchical cluster analysis and Principal Component Analysis.

The resulting visualizations and groupings reveal clusters that align with the discipline assignments for the journals for four of the six disciplines, but also greater overlaps among some journals for two of the disciplines or categorizations that do not necessarily align with their assigned disciplines.

Some journals categorized into a single given discipline were found to be more closely aligned with other disciplines and some journals assigned to multiple disciplines more closely aligned with only one of the assigned disciplines.

The proposed method offers a complementary way to more traditional methods such as journal co-citation analysis to compare journal similarity using data that are readily available through Web of Science.

Aspects of scholarly communication may be investigated from different levels of granularity to reveal and better understand relationships between researchers, research groups, institutions, regions/nations, specializations/disciplines, publications or publication outlets.

Connections that exist between sources of interest may take the form of direct citations, co-citations, co-authorship, co-occurrence of words or subjects, or more recently, latent topics.

Journal similarity comparison has been frequently studied using co-citations. Journal co-citation studies have been carried out on a number of fields including economics (McCain, 1991), information retrieval (Ding, Chowdhury & Foo, 2000), information systems (Marion, Wilson & Davis, 2005), medical informatics (Morris & McCain, 1998), neural networks (McCain, 1998), and semiconductor research (Tsay, Xu & Wu, 2003). 

One reason co-citation studies tend to focus on individual fields is that the journal–journal co-citation matrix that emerges when multiple disciplines are employed can be quite sparse (Boyack, Klavans, & Börner, 2005).

Co-citation data can also be labor-intensive to extract and are not easily available through citation database sources such Thomson Reuters Web of Science (WoS) without downloading all references from a corpus of articles.

Citation-based data may also be used to identify disciplinary or specialization affiliations for journals. This is particularly important for informetrics studies, where the misclassification of journals may affect the ranking of journals within a given field.

Pudovkin and Garfield (2002) developed a journal relatedness factor based on citing and cited journals. The goal of their proposed method was to help identify thematically related journals.

Similarly, Glänzel and Schubert (2003) developed a three-step process for the categorization of journals that involved pre-defined categories, journal classification and article classification for articles in journals with ambiguous subject assignments based on references.

More recently, Rafols and Leydesdorff (2009) compared the outcomes of two algorithms for the decomposition of large matrices against Web of Science Subject Categories and Glänzel and Schubert’s categorization. The four methods they used resulted in similar map outcomes on a large scale. Leydesdorff and Rafols (2009) also investigated the relationships among 170+ Web of Science Subject Categories using a citation matrix consisting of the subject category citation frequencies. They concluded that a classification scheme could be developed using analytical arguments.

Similarly, Leydesdorff and Schank (2008) visualized and animated the disciplinary ties of three seed journals over time to demonstrate relationships among journals and their interdisciplinarity.

Co-citation analysis relies on citing articles to identify the strength of relationships between the units of interest, whether authors, papers or journals; however, it does not consider any attributes of the source of the citations – only that the citations or co-citations exist. Authors such as White (2001) and Ajiferuke, Lu, and Wolfram (2010) have called for a shift in the focus of citation-based research away from citation counts received by an author of interest to the origin of the citation and its characteristics to assess author impact from a different perspective.

This research investigates the use of data derived from citing journals to assess the similarity of cited journals.

The journal citation image of a target journal, which is determined by the list of journals that cite the target journal, provides an indicator of the reach of a journal. When combined with the frequencies of citation by the citing journals, the frequency distribution of citations provided by the citing journals creates a “signature” for each cited journal. These signatures may be compared using various analytical methods.

One possible challenge associated with using the citing journals themselves to create a signature for a cited journal is the potentially high number of citing journals that an influential and prolific journal might attract.

Wang and Wolfram (forthcoming) proposed a method to reduce the computational overhead associated with the citing journal data. Their method of citing discipline analysis uses the subjects/disciplines assigned to the citing journal and the resulting citation frequencies of the citing disciplines to constitute the cited journal’s signature.

Wang and Wolfram (forthcoming) employed citing discipline analysis to explore journal similarity among 40 high impact journals in Information Science and Library Science (ISLS) as classified in Journal Citation Reports (JCR). They found that some of the journals classified into the ISLS category did not map in close proximity to one another based on multidimensional scaling and cluster analysis. A number of the journals included were also classified into allied fields, but did not cluster or appear in close proximity to a number of journals only classified in ISLS.

The authors noted that how journals are classified can impact journal rankings within a given field, where journals from related, but larger, fields may have higher journal impact factors (IF), which can reduce the rank of journals that are directly in the field. They observed that many of the high impact journals were from allied areas to ISLS.

One limitation of their exploratory study was the focus on a single discipline. Could similar affinities or differences in journal similarity be revealed using citing discipline analysis with journals from multiple fields? Also, by looking at multiple fields, does this pull journals also classified into other disciplines further out of an assigned category when included?

The present research is guided by the following questions:
1. To what extent are high impact journals from allied disciplines similar to one another based on the discipline of the articles that cite a given journal?
2. Do journals classified into multiple disciplines more closely align with one discipline than another or serve as bridges between the disciplines to which they are mapped based on the citing discipline distribution?

The field of Information Science and Library Science was selected as the seed discipline based on its interdisciplinary nature and familiarity to the authors. The top 20 journals based on 2012 JCR impact factors were selected. Four additional allied WoS disciplines were also selected based on the co-classification of journals appearing in the top 20 ISLS list with other disciplines, and the affiliation of information science and library science academic units with other disciplinary units, which demonstrates another type of alliance.

The four JCR disciplines selected comprised:
◦ Communication (COMM) – based on the existence of schools of communication & information.
◦ Computer Science, Information Systems (CSIS) – based on a number of iSchools and journal overlap in JCR.
◦ Education & Educational Research (EDER) – based on a number of ISLS units affiliated with colleges/schools of education.
◦ Management (MGMT) – based on the overlap of journals, particularly in Management Information Systems (MIS).

A sixth, more intellectually distant, discipline, namely Geology (GEOL), was also included. Geology was selected based on the outcomes of the UCSD Map of Science (Börner et al., 2012), where Earth Sciences were mapped as distant from the Social Sciences. By including journals from a more distant discipline, the ability for the citing discipline method to distinguish between more closely aligned and distant disciplines could be tested, where the distinctiveness of allied disciplines may be less defined by including a more distant discipline in the analysis.

A total of 120 journals were studied over the time period 1987–2012. To allow for a comparison over time, the journals were subdivided into three time periods: 1987–1995, 1996–2004, and 2005–2012. 

The data collection method for determining the frequency distribution of citing disciplines used in Wang and Wolfram(forthcoming) was adopted for the present study.

The “Create Citation Report” option in WoS was selected to identify all citing articles. The number of Citing Articles was then selected to retrieve the list of citing articles. The WoS “Analyze Results” feature was next selected for the list of citing articles. On the Results Analysis page, “Research Areas” were selected as the ranking field to provide the tabulated list of citing disciplines.

Salton’s Cosine measure was used determine the similarity between pairs of journals, resulting in a symmetric similarity matrix (Ahlgren, Jarneving, & Rousseau, 2003; Egghe & Leydesdorff, 2009; Leydesdorff, 2006).

Multidimensional scaling (MDS) analysis and hierarchical cluster analysis using SPSS v.20 were employed to visualize and categorize the relationships among the journals for each time period.

To provide a complementary analysis of the hidden groups that may be present in the data, SPSS’s Factor Analysis using Principal Component extraction with varimax rotation was also applied to the data using routines.


Fig. 1 shows the MDS locus of 70 selected journals in the first time period (1987–1995). The raw stress value was 0.01294,and the stress-I was 0.11376.

Only five clusters are shown because a distinctive sixth cluster did not emerge for this time period.

At the five-cluster level of assignment, journals from COMM, EDER, and MGMT form coherent clusters, although based on the MDS map some journals in each field are more closely located to journals in an allied discipline.

The fourth cluster combines journals from ISLS and CSIS. It is possible that the relatively small number of purely CSIS journals for this time period did not provide enough data for these journals to cluster into separate groups.

The MIS journals are situated in the ISLS cluster, but are located between the Library and Information Science (LIS) journals and MGMT journals.

The fifth cluster on the right side of the map consists of GEOL journals and, as would be expected, is quite distinctive from the other five disciplines.

The location of some journals on the map suggests that they are more similar to journals in one of the other given categories in JCR. The Journal of Advertising Research (JAR), for example, is situated with management journals but is only classified with COMM (and with Business, but this discipline is not included in this study).

Some journals classified in two disciplines served as a bridge between the two disciplines in the map. For instance, Science Communication (SCOMM), although classified with COMM journals, is situated in the ISLS cluster, but is in closer proximity to the COMM journals. In the remaining two time periods, SCOMM clusters with the COMM journals. Social Science Computer Review (SSCR), which is classified with ISLS and clusters with the discipline, is situated between ISLS, COMM and/or MGMT journals in each time period. The same is observed with Telecommunications Policy (TP), which bridges ISLS and MGMT for each of the periods of study.

The outcome for the second time period (1996–2004) appears in Fig. 2. The raw stress and stress-I values are 0.01648 and 0.12798, respectively.

In this map, 93 journals were categorized into six clusters, with CSIS separating from the ISLS cluster during this time period.

Of note is the greater number of journals assigned to one or more disciplines but aligning more closely to another discipline or only one of the assigned disciplines.

As an example, the Academy of Management Learning& Education (AMLE) and JAR are situated in the management cluster, and are located relatively far from their assigned disciplines, EDER and COMM, respectively.

Decision Support Systems (DSS) clusters with the ISLS journals but is classified in CSIS only. The same classification is observed for this journal in the third time period.

The Journal of the American Medical Informatics Association (JAMIA) is classified in ISLS and CSIS, but clusters with the CSIS journals for the remaining time periods.

Similarly, the Journal of Health Communication (JHC), which is classified in COMM and ISLS, does not appear to be similar to other ISLS journals and is situated more closely to COMM journals and clusters with them.

The International Journal of Geographical Information Science (IJGIS), which is classified with ISLS journals clusters with CSIS journals for this time period and the third period, although it is at the periphery of the cluster, perhaps indicating the CSIS discipline is the best match of the disciplines studied, but is not a very close match. Proximally, it is situated between CSIS and GEOL journals, which may indicate at least a peripheral similarity to some GEOL journals.

Again, the GEOL journals all cluster together farther from the other disciplinary groups.
Results for the third time period (2005–2012) appear in Fig. 3. The six clusters roughly correspond to the six disciplines. The results of raw stress calculation and the stress-I calculation are still relatively low, at 0.02224 and 0.14914, respectively.

Additional journals classified in COMM map closely to and cluster more closely with MGMT journals.

The International Journal of Computer-supported Collaborative Learning (IJCSCL) is classified with both EDER and ISLS but is clearly situated and clusters with the EDER journals.

Business Strategy and the Environment (BSE), although clustered with MGMT journals, appears to be pulled toward the GEOL journals, indicating a possible relationship with some of these journals.

Once again, the GEOL journals are distinctly clustered away from the remaining journals.

With each time period, more journals cross disciplinary boundaries by clustering with journals from allied disciplines or by mapping more closely to journals in allied disciplines. There may be several influencing factors to account for this observation.

First, the journals indeed may be becoming more interdisciplinary in their publication coverage, thereby attracting more citations from publications in other disciplines.

Second, the journals themselves may not be more interdisciplinary in their coverage, but are now more easily discovered by authors in other disciplines given the wider availability of federated search tools.

Third, with a greater number of journals included in the analysis for each time period, the distinctiveness of the citing discipline signatures may be decreasing, so some journals classified in allied disciplines may appear more similar to one another.

To determine common dimensions from the dataset, a Principal Component Analysis was conducted in SPSS for each time period. Outcomes for the Kaiser–Meyer–Olkin measure of sample adequacy (above 0.7) and Bartlett’s Test of Sphericity (p < .05) indicate the data were appropriate for PCA for all three time periods.


In this period, the six components explain 88.1% of the total variance that correspond to the six disciplines.

There are five journals underlined in the Table 2 that belong to two components, introducing inter-factorial complexity (Van den Besselaar & Heimeriks, 2001; Leydesdorff, 2007), including three MGMT journals (ISJ, ISR, MISQ).

SCOMM and DSS are assigned only to COMM and CSIS by WoS, respectively, but they also appear in the ISLS component, which supports the MDS and clustering outcomes. DSS continues to also load with ISLS for the remaining time periods.

Similarly, JAR and AMLE, journals classified by WoS only in COMM and EDER, respectively, load with the MGMT component for all the time periods in which they appear, but not their classified discipline, lending support for the re-classification of these journals.

The ISLS journals ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS, and MISQ are also classified at CSIS, but do not load with the CSIS component.

Outcomes for 1996–2004 appear in Table 3.

As with the first time period, several MIS journals (IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ) load to both MGMT and ISLS.

SSCR, which is classified with ISLS only, loads into the ISLS component, but also loads with a higher value into the COMM component, perhaps indicating the need for an additional classification assignment.

IJGIS does not load into any of the six components for second or third time period, lending evidence to the peripheral nature of the journal to the six fields studied and indicating it might be misclassified in ISLS.

Component outcomes for 2005–2012 appear in Table 4.

Similar to the previous time period, a growing number of MIS journals load to the ISLS and MGMT components.

As with IJGIS, two other journals, Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM) do not load into any of the six components, indicating a poor association with the six disciplines.

The boundaries between ISLS and CSIS are not as clear in the MDS and cluster analysis outcomes, where combinations of computer science, library and information science and management information systems journals may cluster together depending on the time period. These results may be influenced by the fact that a number of journals in the ISLS area are also categorized in the CSIS or MGMT category, thereby strengthening their relationships.

Despite the influence of the assigned discipline(s) – which then strengthens the disciplinary relationship(s) through journal self-citations, whether or not it is the best fit – some journals appear to be misclassified based on the disciplinary designations of the citations they attract.

As a prime example, JAR, which is only assigned to the COMM field, is situated in the MGMT category for each of the time periods studied for the MDS and clustering analysis as well as with the Principal Component Analysis.

A similar outcome is observed for AMLE for the two time periods in which it is included. It is classified as an EDER journal, but is situated with the MGMT journals based on the analyses conducted.

Several other journals are classified in more than one discipline, but clearly associate only with journals in one of the disciplines. JHC is classified in COMM and ISLS but clusters only with COMM journals. The same is observed for IJCSCL, which is classified in ISLS and EDER but groups only with EDER journals for each of the grouping methods used.

Other journals appear to move between disciplines over time. SCOMM is classified in COMM, but the MDS and clustering outcome places it initially with the ISLS journals for the first time period, and then in the COMM cluster and further away from ISLS in last two time periods.

Some journals, such as JCMC and SSCR, appear to be situated near borders between disciplines, which point to their interdisciplinary appeal or may indicate they serve as bridges between the disciplines.

A number of the ISLS journals that are considered to be library and information science journals (Nisonger & Davis, 2005) appear between CSIS journals and those in the MGMT cluster. In particular, MIS journals appear between MGMT and the ISLS/CSIS cluster for the first time period.

One application of citing discipline analysis that emerges from this analysis is that of decision support for the additional assignment or reassignment of journals to one or more disciplines.

Journal disciplinary classifications should be revisited over time to accommodate shifts in how journals are being cited by other disciplines. ... However, the shifts are at least an indication that the subject affiliations of the citing the journals are changing.

Citing discipline analysis provides, with some modest programming, a relatively easily implemented method for assessing the similarity of journals within disciplines or across allied disciplines that is computationally less expensive than using citing journal-based data.

The analyses reveal distinct groupings of journals based on their disciplinary assignments. As observed earlier in Wang and Wolfram (forthcoming), who only examined journals in a single field, the current research has demonstrated that citing discipline analysis can provide coherent and meaningful disciplinary groupings for journals in allied fields, even when journals from a more intellectually distant field are included.

The clustering and proximity of some journals classified in allied fields has changed over time, perhaps indicating a changing citing relationship between these fields.

2014年6月21日 星期六

Jensen, P., & Lutkouskaya, K. (2014). The many dimensions of laboratories’ interdisciplinarity. Scientometrics, 98(1), 619-631.

Jensen, P., & Lutkouskaya, K. (2014). The many dimensions of laboratories’ interdisciplinarity. Scientometrics, 98(1), 619-631.

Scientometrics

本研究提出六種指標來測量研究機構的跨學科性。最廣義的來說,跨學科性可以視為不同學科某種程度的整合 (Weingart and Stehr 2000; Porter and Rafols 2009; Marcovich and Shinn 2011; Wagner et al. 2011; Rafols et al. 2012),為了將這個想法轉換為量化的指標,本研究認為需要考慮三個問題:
1. 如何定義一個學科
2. 在什麼層次達到整合
3. 學科連結需要到達什麼程度

在學科的定義上,本研究提出三種方式:一、因為是分析CNRS的實驗室,自然可採用CNRS的學科組織(disciplinary organization),包括10個研究所(institutes)以及進一步細分成的40個組(sections);二、如同其他先前的研究,使用WoS(Web of Science)的224種期刊主題分類(Journal Subject Categories, JSCs);三、將文件根據共同的參考文獻,以叢集演算法(clustering algorithms)由下往上地(bottom-up)歸類成認知叢集(cognitive clusters)。在整合的層次,則探討實驗室與論文兩個層級。

以實驗室的跨領域程度來說,較簡單的方式可以定義為:
此處的pi是實驗室的論文在期刊主題分類JSC i上的比例。

除了上述的定義之外,本研究還使用的Stirling’s (2007)方法來表現多樣性的三個不同面向:不同類別的數量(variety)、在各類別上的分布均勻程度(balance)、以及表現類別間的差異(disparity) (Porter and Rafols 2009):
此處的是主題分類JSC i 和 JSC j的相似性,並且此一相似性以cosine測量主題分類間的引用情形得到。(Porter and Rafols 2009).

為了進一步了解實驗室的跨領域多樣性是否在單一論文的認知層次達成,如同上述的情形,計算單一論文的跨領域多樣性時,可以利用下面的方式:
此處的pai是此論文引用的參考文獻在期刊主題分類JSC i上的比例。進一步用實驗室發表的論文考慮實驗室的跨領域程度時,可以將所有論文的跨領域多樣性加以平均,如

此處的 #pap 是實驗室發表的論文數量。

此外,另兩種指標分別是主流引用外的主題分類比例以及不同機構的人員合作占論文全體比率,分別如下所示:


最後一種指標,先以書目耦合(bibliographic coupling) (Kessler 1963)產生論文之間的關連,計算方式如下:
此處的where #common_refsij 是論文 i 和 j 共同引用的參考文獻數量, #refsi 和 #refsj 分別是論文 i 和 j 包含的參考文獻數量。接下來以書目耦合關連建立論文網路,希望在網路上引用文獻相似的論文會聚集形成叢集。因此,接下來Blondel et al. (2008)的演算法,劃分網路成論文的叢集。整個方法可參見Grauwin and Jensen (2011),結果共劃分成250個叢集。然後以下面的方式計算實驗室在認知叢集上的多樣性
此處的 p_i 和 p_j 分別是實驗室的論文屬於叢集 i 和 j 的比例。

以六種指標計算每一個實驗室的跨學科多樣性後,接下來以主成分分析(Principal Component Analysis, PCA)進行分析,四個主要的成分分別是
1) 實驗室在各種多樣性指標的綜合表現
2) 實驗室連結的學科的認知距離(cognitive distance)
3) 實驗室在實驗室層級或論文層級具有跨學科性
4) 論文發表的期刊具有跨學科的主題分類或是與其他不同機構的實驗室合作。

Interdisciplinarity is as trendy as it is difficult to define. Instead of trying to capture a multidimensional object with a single indicator, we propose six indicators, combining three different operationalizations of a discipline, two levels (article or laboratory) of integration of these disciplines and two measures of interdisciplinary diversity.

Interdisciplinarity means, at the most generic level, some degree of integration of different disciplines (Weingart and Stehr 2000; Porter and Rafols 2009; Marcovich and Shinn 2011; Wagner et al. 2011; Rafols et al. 2012).

To transform this idea into quantitative indicators, we need to answer three questions:
1. How to define a discipline?
2. At what level the integration is achieved?
3. What is the degree of disciplinary linkage achieved?

There are several ways to define a discipline from a scientometrics’ point of view. Since we are dealing with CNRS labs, the most natural would seem to use the disciplinary organization of CNRS in 10 ‘‘institutes’’ and 40 subdisciplinary ‘‘sections’’. A convenient alternative is to use the 224 Journal Subject Categories (JSCs) used by Web of Science (WoS). Finally, instead of using institutionally predefined divisions of science, one could use a more bottom-up definition of ‘‘cognitive clusters’’. To obtain these clusters, we use the roughly 300,000 French articles published between 2007 and 2010 and group them into ‘‘cognitive clusters’’ using clustering algorithms based on shared references.

In this paper, we will use three definitions of ‘‘discipline’’ and two integration levels (laboratory and article) to calculate six partial interdisciplinary indicators.

We adopt Stirling’s (2007) approach to capture the different facets of diversity : ‘variety’, ‘balance’ and ‘disparity’.

‘Variety’ characterizes the number of different categories, ‘balance’ characterizes the evenness of the distribution over these categories and ‘disparity’ characterizes the difference among the categories, usually based on some distance.

A simple indicator of the spread of the disciplines where a laboratory publishes is given by:
where pi is the proportion of articles of the laboratory in JSCi.

As we would like to include the idea of ‘‘distance’’ between disciplines, we calculate the diversity indicator (Stirling 2007; Porter and Rafols 2009) which combines both the spread of the disciplines through the pi and the distance between them.
where sij is the cosine measure of similarity between JSCs i and j. Practically, sij is measured through the citations from publications in JSCsi to publications in JSC j (Porter and Rafols 2009).

To further characterize a lab’s interdisciplinarity, it is useful to introduce an indicator of the interdisciplinarity of single articles, to test whether interdisciplinarity is achieved at this cognitive level.

Specifically, the interdisciplinary diversity of a single article is calculated as:
where pai is the proportion of articles’ references in JSCi.

To quantify the interdisciplinarity of the papers published by a lab, we aggregate the articles’ diversity indicator art_div_corr at the laboratory level by averaging over all the articles published by that laboratory:

where #pap is the number of articles of the lab for which at least one reference was identified.

Then, we choose a threshold to define the most common JSCs for each institute. ... We therefore choose a threshold value of 90 %. ... Then, for each laboratory, we count the percentage of articles outside this 90 % list and normalize by the expected value, i.e. the average value 0.1.


whereare the frequencies of the JSCs that do not belong to the Institute’s JSC main list.

Interdisciplinary collaborations can also be detected by copublications between scientists belonging to different CNRS Institutes. We compute a fifth indicator by calculating the proportion of a lab’s publications that involve authors from other Institutes

where the sum counts the number of articles of the lab involving at least two institutes and
#articles is the total number of articles published by the laboratory.

To build these ‘‘cognitive disciplines’’, we use bibliographic coupling (BC) (Kessler 1963) between the 300,000 papers published by French laboratories in the period 2007–2010 and compiled by the WoS.
where #common_refsij is the number of common references for articles i and j, and #refsi,
#refsj are the numbers of references of articles i and j, respectively.

In comparison to a co-citation link (which is the usual measure of articles’ similarity), BC offers two advantages: it allows to map recent papers (which have not yet been cited) and it deals with all published papers (whether cited or not).

This reinforcement facilitates the partition of the network into meaningful groups of cohesive articles, or clusters. A widely used criterion to measure the quality of a partition is the modularity function (Fortunato and Barthe´lemy 2007), which is roughly is the number of edges ‘inside clusters’ (as opposed to ‘between clusters’), minus the expected number of such edges if the partition were randomly produced. We compute the graph partition using the efficient heuristic algorithm presented in (Blondel et al. 2008). The whole method is described in (Grauwin and Jensen 2011).

Applying this algorithm yields in a partition of French papers into roughly 250 clusters containing more than 100 papers each.
where p_i is and p_j are the proportions of the labs’ papers belonging to clusters i and j respectively.

On average, articles refer to papers from almost 10 different disciplines (9.8 JSC). .... However, when considering those JSC that are used in more than 10 % of the reference list, this average drops to 2.7. This means that, on average, an article spreads its references on 3 main JSCs and 7 additional which benefit from roughly a single reference.

An average laboratory publishes in journals belonging to 34 different JSCs ...

PCA1: combined interdisciplinarity The main axis represents a combination of the various interdisciplinarity indicators.

PCA2: short or long cognitive distance This axis distinguishes those labs that connect distant or nearby disciplines.

PCA3: article or laboratory interdisciplinarity This axis distinguishes labs that achieve interdisciplinarity either at the laboratory or article level.

PCA4: diversity of publications’ JSCs or diversity of collaborations This axis distinguishes labs that publish in journals belonging to different JSCs (high lab_jsc_bal) from labs that co-publish with labs from different CNRS Institutes (high lab_inst_cop_bal).

We have computed the six indicators for the 680 laboratories which have published more than 50 papers over 2007–2010. To allow comparisons and statistical analysis, since the absolute values have no intrinsic meaning, we have scaled all the values to achieve an average value of 0 and a variance of 1. We then carried out a principal component analysis of the (680 9 6) matrix using the free software R (www.r-project.org/). More precisely, we used prcomp from the ‘stats’ package, without any axes rotation.

First, let us note that using the first four PCA axes gives an overall view about the interdisciplinarity practices of each lab. This view has been compared to expert knowledge, namely scientists working in those labs or scientific advisors from CNRS. This comparison, carried out for about 20 different labs from all the disciplines, suggests that these indicators characterize interdisciplinarity
in a meaningful way.

A major drawback of our method is that we cannot distinguish real interdisciplinary collaborations, giving rise to new concepts or to a coherent new scientific field, from simple pluridisciplinary practices that merely juxtapose different disciplines, as when historians use characterizing tools from physics. It seems difficult to learn much about the cognitive dimensions of interdisciplinarity from an automatic analysis of metadata of the papers.

2014年6月20日 星期五

Porter, A. L., & Rafols, I. (2009). Is science becoming more interdisciplinary? Measuring and mapping six research fields over time. Scientometrics, 81(3), 719-745.

Porter, A. L., & Rafols, I. (2009). Is science becoming more interdisciplinary? Measuring and mapping six research fields over time. Scientometrics, 81(3), 719-745.

Scientometrics

本研究將跨學科研究(interdisciplinary  research)操作化的定義為:由團隊或個人從兩個或以上的知識體系(bodies of knowledge)或研究實務整合它們的觀點/概念/理論、工具/技術以及資訊/資料的一種研究模式,也就是這類研究其知識來源具有多樣性,然後分析六個研究領域在1975年和2005年的跨學科程度變化。跨學科指標的計算以引用期刊在WoS (Web of Science)上的主題分類(Subject Categories, SCs)為基礎,並且配合科學映射圖(science maps)表現科學產出在主題分類上的分散情形(dispersion)。整個分析的流程包含五個步驟:


一、將跨學科性的測量操作化。
二、建構主題分類間的相似性矩陣,做為計算整合性指標之用。
三、對相似性矩陣進行因素分析(factor analysis),將主題分類分群成為巨型學科(macro-disciplines)以便進行視覺化。
四、產生科學映射圖。
五、選取六個主題分類,做為目前的基準與未來探索。


針對操作化跨學科性的測量有幾點必須說明:首先根據Stirling的看法,探索跨學科性時,需要針對引用的學科數量、引用在學科間的分布情形、類別的相似性等面向進行研究[RAFOLS & MEYER, FORTHCOMING]。其次,本研究認為知識整合是一種認知範疇(an epistemic category),因此跨學科性指標應該建立在研究結果的內容,而不是團隊的成員,部門組織或合作上。最後,跨學科性的測量通常以引用文獻的期刊所屬的主題分類為基礎,但書目計量學的研究社群已經提出主題分類有一些問題,例如期刊叢集的研究指出僅有約50%的叢集結果和主題分類相近[BOYACK & AL., 2005; (BOYACK, personal communication, 14 September 2008)],根據引用網路得到的分類結果和主題分類之間也沒有很好的符合[LEYDESDORFF, 2006, P. 611]。但這些結果僅對科學映射圖產生有限度的影響,並且在測量整合性上,主題分類目前還是最被廣泛使用的分類資源。

本研究用來測量整合性指標[RAFOLS & MEYER, FORTHCOMING]的公式,由Rao-Stirling提出的多樣性測量方式 [STIRLING, 2007],如下:
此處pi是給定的論文上引用的參考文獻來自主題分類 i 的比例,sij是主題分類 i 和 j 的相似程度,利用cosine測量。由於許多研究 例如[GRUPP, 1990; HAMILTON & AL., 2005, or ADAMS & AL., 2007]都以Shannon或Herfindhal提出的方式測量整合性,Shannon的多樣性測量方式如
Herfindhal的多樣性測量方式如


但這兩種方法都未考慮類別間的不同;反之Rao-Stirling的多樣性則同時考慮類別數量多寡、類別上的分布平衡和類別間的相似性等三個方面。因此本研究比較此一整合性指標與Shannon和Herfindhal多樣性。

本研究以主題分類被引用的次數為資料,對每一對主題分類進行cosine測量這兩個主題分類間的相似性。當兩個主題分類被大部分的論文共同引用時,它們之間便會有很高的相似度;反之,兩個主題分類共同被引用的情形很稀少時,cosine的值接近於零。完成相似性矩陣的建立後,以主成分分析(Principal Components Analysis, PCA)進行因素分析,以最大變異量轉換(Varimax rotation)產生20個因素,將每一主題分類以其具有最高負荷的因素進行歸類。每一個因素對應一個巨型學科,某些無法歸類的主類分類另外歸於一個巨型學科,結果共有21個巨型學科。

然後以主題分類在21個因素上的負荷值為特徵,再以cosine測量主題分類之間的相似性。以Pajek將主題分類之間的相似性映射成網路圖,過濾相似性在0.6以下的連結線,做為科學映射圖。科學映射圖上呈現每一個主題分類、相對的重要性、以及彼此間的關連程度,目的在於在巨型學科間找出特定研究的主體,發現相互關連在時間上的變化以及主要的跨學科關連,更重要的是發現做為知識來源的期刊是來自於密切關連的學科或是跨越完全不同的領域。


本研究選取生物科技與應用微生物學(Biotechnology & Applied Microbiology)、電子電機工程(Engineering, Electrical & Electronic)、數學(Mathematic)、醫學(Medicine – Research & Experimental)、神經科學(Neurosciences)、物理(Physics – Atomic, Molecular & Chemical)等六個主題分類。

研究結果發現:30年間論文的平均作者數、平均參考文獻數和引用的學科數量都有很大幅度的增加,但是從跨學科指標的增加並不大。造成上述現象,可能是由於雖然引用的主題分類數量有明顯的增加,但每篇論文平均引用的參考文獻數量增加地更快,使得在不同主題上的引用比例的實際改變變得不如預期中的重要;另外,許多主題分類的引用較傾向於鄰近的主題分類,但是鄰近區域的主題分類有較高的相似值,對於多樣性的貢獻較低;最後是某些較跨學科研究的領域其測量的整合性已經到達飽和了。從科學映射圖的結果也指出論文引用的分布仍然主要集中於某些鄰近的學科領域。此外,本研究也發現Rao-Stirling的多樣性測量與Herfindhal和Shannon的測量都有很高的相關性,分別為0.91(標準差0.07)及0.88(標準差0.07) 。

Here we investigate how  the degree of interdisciplinarity has changed between 1975 and 2005 over six research domains. ... The results attest to notable changes in research practices over this 30 year period, namely major increases in number of cited disciplines and references per article (both show about 50% growth), and co-authors per article (about 75% growth). However, the new index of 
interdisciplinarity only shows a modest increase (mostly around 5% growth). Science maps hint 
that this is because the distribution of citations of an article remains mainly within neighboring 
disciplinary areas.

We measure how integrative particular research articles are  based on the association of the journals they cite to corresponding Subject Categories  (“SCs”) of the Web of Science (“WoS”)

And, we present a practical way to map  scientific outputs, again based on dispersion across SCs.

This report operationally defined interdisciplinary  research as: 
x a mode of research by teams or individuals that integrates 
x perspectives/concepts/theories and/or 
x tools/techniques and/or 
x information/data 
x from two or more bodies of knowledge or research practice. 

Our approach here is to investigate changes of degree of interdisciplinarity over time  using various established indicators (e.g. number of disciplines cited, percentage of  citations within-field), together with a new indicator developed the NAKFI evaluation  team [PORTER & AL., 2007]: 
Integration – reflecting the diversity of knowledge sources, as shown by the breadth  of references cited by a paper. 

Following Stirling’s heuristic, we have previously argued that in order to explore interdisciplinarity, one needs to investigate multiple aspects, namely: the number of disciplines cited (variety), the distribution of citations among disciplines (balance), and, crucially, how similar or dissimilar these categories are (disparity) [RAFOLS & MEYER, FORTHCOMING]. 

The computation and visualization of the interdisciplinarity measure has taken five  steps, presented consecutively in this section: 
1. Operationalization of an interdisciplinary measure (the Integration index or disciplinary diversity)
2. Construction of a similarity matrix among Subject Categories that is used to compute the Integration index
3. Grouping via factor analysis of the SCs into macro-disciplines using the similarity matrix as a base to facilitate visualization
4. Generating science maps
5. Selection of a bibliometric sample of 6 SCs, to serve as benchmarks here and in future explorations. 

In other words, since knowledge integration is an epistemic category, indicators of interdisciplinarity should be based on the content of the research outcomes rather than on team membership, departmental affiliations, or collaborations (see illustrations in case studies in RAFOLS & MEYER, 2007). 

The bibliometric community has noted that the SCs have some problems. In journal clustering exercises, only about 50% of clusters were found to be closely aligned with SCs [BOYACK & AL., 2005; (BOYACK, personal communication, 14 September 2008)]. Poor matching between SCs and classifications derived from citation networks has also been reported [LEYDESDORFF,
2006, P. 611], but surprisingly the mismatch only has limited effect on the corresponding science maps [RAFOLS & LEYDESDORFF, UNDER REVIEW].

Nonetheless, the SCs offer the most widely available categorization resource that we could ascertain for the purpose of providing an accessible measure of Integration.

As derived in RAFOLS & MEYER [forthcoming], the formula for the Integration index can be expressed as:

where pi is the proportion of references citing the SC i in a given paper. The summation is taken over the cells of the SC x SC matrix. sij is the cosine measure of similarity between SCs i and j (the cosine measure may be understood as a variation of correlation). Here this matrix sij is based on a US national co-citation sample of 30,261 papers from Web of Science as explained below in detail. 

This Integration measure (aka, Rao-Stirling’s diversity) can be compared with Shannon diversity: 

or with Herfindhal’s diversity (the complement of Herfindahl’s concentration):

The power of the Integration index is that it characterizes interdisciplinarity in terms of the diversity of knowledge sources of papers, using a general formulation of diversity [STIRLING, 2007] rather than an ad hoc indicator.

A number of researchers have used these traditional measures of diversity, such as Shannon or Herfindhal, to measure interdisciplinarity [E.G. GRUPP, 1990; HAMILTON & AL., 2005, or ADAMS & AL., 2007]. These measures do not take into account how different the categories are, whereas our Integration measure reveals increased diversity only when added categories are significantly different.

In particular, a broad national sample of articles from WoS is used to create the sij matrix that underlies the metrics used for computing Integration. First we describe the sample used as a basis for the similarity matrix; second, the construction of the matrix.

We combine six separate weeks of all papers in WoS, with one or more authors having a USA address, sampled during 2005–2007, to obtain 30261 articles. This provides a broadly based, yet manageable base sample. We processed the “Cited References” of these abstract records to identify the “Cited SCs.”

Our sample of 30261 WoS articles contains 1,020,528 cited references (an average of 33.7 per article). Of those, our thesauri link 768,440 to a particular Subject Category. Another 28,000 have been checked and assigned to “not being in an SC.”

For our purposes in addressing cited SCs, the list includes a few more than the current set, for a total of 244 SCs. The sample contains 1,114,930 instances of cited SCs.

The 30261 articles, by 244 SCs, described allow for construction of a co-citation similarity matrix, sij, using Salton cosine [SALTON & MCGILL, 1983; AHLGREN & AL., 2003].

The values of sij are high (i.e. closer to one) when SCs i and j are co-cited by a high proportion of articles that cite one or the other. The cosine value approaches zero when two SCs are rarely cited together.

For various purposes and in particular for visualization, it helps to consolidate the narrow research areas of the ISI SCs into larger categories, which we call “macro-disciplines.”

We base our grouping of SCs on a type of factor analysis – Principal Components Analysis (PCA) – following a similar methodology to that developed by LEYDESDORFF & RAFOLS [2008] to cluster SCs into macro-disciplines.

Within VantagePoint, we constructed the matrix of cosine similarities for the 244 cited SCs by 244 cited SCs described in the previous section. ... We explored various factor analysis solutions, eventually adopting a 20-factor solution (Varimax rotation). ... The 21 macro-disciplines reflect this factor solution.

So, to a considerable degree, named sub-disciplines do not fully coalesce within a single macro-discipline. This warns that the evolving research enterprise does not neatly conform to the traditional scholarly disciplines.

These maps present the SCs, their relative importance in size, and how related they are to each other over all science. The main aim of these science maps is to locate particular bodies of research among the macro-disciplines. ... That can help identify changes in degree of interrelationship over time, and key cross-“disciplinary” relationships that might benefit from nurturing. It should also be informative to see whether knowledge sources of a set of publications are coming from research domains that are closely related (little interdisciplinarity) or that span very disparate domains (high interdisciplinarity). 

We then construct a new Salton cosine similarity matrix among SCs using the loadings of each SC on the 21 factors (as discussed in the previous subsection). This matrix is then uploaded into the network analysis software Pajek [BATAGELJ & MVAR, 2008]. In Pajek, the minimum similarity threshold was arbitrarily set to 0.6 (this choice was found to provide a good readability-to-accuracy trade-off) and the SCs were distributed in a 2-D plane according to their similarities, to obtain a base science map.

Since research collaboration is often (and sometimes mistakenly) associated with interdisciplinarity, we examine measures of co-authorship. ... However, within research domain, the number of authors per paper has escalated remarkably, with about 75% average growth. This increase ranges from 48% in Math and 54% in Physics-AMC to 90% in Neurosciences. 

Before turning to Integration scores, we consider the number of distinct SCs that one article cites. ... Table 2 and Figure 4 show a sturdy increase in the breadth of citing in all six of these research domains (about a 50% growth on average). 

Integration scores are tabulated in Table 2 and shown in Figure 5. We see that over time, there is a modest increase in Integration scores and that math researchers are notably less integrative in their citing patterns. However, math has the highest relative growth (39%) whereas other SCs’ growth ranges from 3% to 14% (5% on average). t-tests between the 1975 and 2005 samples show these differences to be highly significant (<.005 for EE, assuming either equal or unequal variances; all others even more highly significant).

Pearson’s correlation between Integration and Herfindhal takes a mean value of 0.91 (standard deviation = 0.07) and between Integration and Shannon, a mean value of 0.88 (standard deviation = 0.07). These high correlations confirm that Integration is very closely associated with traditional diversity indicators – as could be expected by construction.

The main finding is that Integration scores increase over time, but significantly less so than other indicators, such as percentage of single-authored papers, mean authors per paper, and mean number of disciplines per paper.

First, although the number of cited SCs increases significantly, since the average number of references in a paper also shows a quicker increase (see central columns in Table 2), the actual change in the proportions of citation to different SCs is not as important as could be expected.

Second, as we will show in Figures 7 through 10, the citation patterns of a given SC tend to be with SCs in its vicinity. Since these neighboring SCs have high similarity values with the one investigated, their contribution to Integration (to diversity) is smaller than in other indicators. This means that the Integration score “deflates” the diversity recorded by Shannon or Herfindahl because most of the cited SCs are not very different from the SC doing the citing.

This is much easier to convey using science maps that directly show the three aspects of disciplinary diversity, namely:
1. the variety of “disciplines” (i.e., discrete research areas, the SCs, shown by the number of nodes in the map)
2. the balance, or distribution, of disciplines (relative size of nodes)
3. the disparity, or degree of difference, between the disciplines (distance between the nodes)

These maps were created followed the techniques developed in LEYDESDORFF & RAFOLS [2008], in the context of the current interest in science mapping [MOYA-ANEGON & AL., 2004; BOYACK & AL., 2005; MOYA-ANEGON & AL., 2007]. ... In the figures presented in this article, we only label groups of SCs on the basis of macro-disciplines found by factor analysis, as explained in the methodology. 

However, the perspective provided by the Integration score and the science maps suggests that the practice of interdisciplinarity in citations occurs mainly between neighboring SCs and has undergone a much more modest increase (on average only 5%, excluding math).

This is mainly for two reasons: first, although the number of cited SCs has increased, the growth of citations means that the increase in the proportion of citations to new SCs is small; second, the newly cited SCs tend to be in the vicinity of the previous ones – hence they don’t add as much interdisciplinarity as they would if they were very disparate/distant disciplines. Moreover, for already very interdisciplinary SCs, such as Neuroscience, the indicator may have a certain “saturation” effect.