顯示具有 Pajek 標籤的文章。 顯示所有文章
顯示具有 Pajek 標籤的文章。 顯示所有文章

2014年8月10日 星期日

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers. Scientometrics, 94(3), 1161-1173.

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers.Scientometrics, 94(3), 1161-1173.

network analysis

本研究以發表場域-作者-耦合(Venue-Author-Coupling,VAC)、期刊共被引分析(journal co-citation analysis)、主題分析(topic analysis)和連結編輯委員會成員(interlocking editorial board membership)等四個面向分析資訊科學與圖書館學的期刊網絡。這個研究分析的期刊範圍為2008年 JCR (Journal Citation Report)資訊科學與圖書館學分類的58種期刊,在2005到2009年間的出版資料。分析資料的相關數據如Table 1:


本研究利用VAC代表期刊的生產者(producers)的相似性,根據每一對期刊間相同的作者數量測量它們的接近程度,其原理建立在作者會選擇主題或社會性相似(thematically or socially similar)的期刊發表。期刊共被引分析(McCain, 1991)計算每一對期刊被共同引用的次數,本研究用來測量作品(artifacts)間的相似程度。本研究以修改自LDA模型(Blei et al. 2003)的ACT(Author-Conference-Topic)模型(Tang et al., 2008)透過關鍵詞(keywords)在主題上的分布以及主題在作者及發表場域(期刊)上的分布,本研究以餘弦(cosine)測量評估期刊之間的相似程度。連結編輯委員會成員則是編輯委員會上的共同成員數測量期刊間的相似程度。兩種期刊間共同的成員愈多,代表這兩種期刊在認知上或是社會性上愈相似。

根據上面的四種期刊間的相似程度所得到的結果,除了進行階層式集群分析(hierarchical cluster analysis)之外,也用來建立網絡,以Kamada-Kawaii 法呈現網絡的型態。分析得到的四種網絡並且以二次指派程序(Quadratic Assignment Procedure) (Lawler 1963)比較網絡之間可能的相關性(correlation)。

VAC方法得到的期刊網絡如下
四個集群分別為MIS(黃)、IS(藍)、LS(綠)以及專門性期刊(紅)。其中的MIS期刊集群與其他的集群相當分離。IS與LS距離較近。相較於其他三個期刊集群,專門性期刊彼此間的連結較弱。
期刊共被引分析所得的網絡如下:

主題模型產生的五個主題如Table 2
五個主題在網絡上的分布如下圖
MIS(黃)仍然與其他集群較為分離,但與健康和傳播(communication)等專門性期刊的距離較近。IS(藍)和LS(粉紅)的位置與VAC和期刊共被引分析的網絡上有所不同。圖書館服務與實務(綠)與專門性期刊和LS很接近。

在利用連結編輯委員會成員的期刊網絡上,有10種期刊沒有和其它期刊有共同編輯委員。其餘的集群分為四群。以傳播研究相關的期刊是新增加的集群(綠)。

四個網絡的QAP結果如Table 3

總結以上,在JCR的資訊科學與圖書館學分類下約略可以將期刊分為四個集群:MIS、IS、LS和傳播相關的期刊。MIS相較來說較為獨立。另外,QAP的結果可以看到編輯委員會成員的結果與期刊共被引分析有很高的相關性,其原因可能是由於擔任編輯委員的研究人員往往有較好的學術成就,被引用的機會較高。編輯委員會成員與VAC有較高的相關性,其原因也可能是編輯委員有較高的生產力。運用多種面向的分析可以較全面地了解整個學術傳播網絡。

Fifty-eight journals from the Information Science and Library Science category in the 2008 Journal Citation Report were studied and the network proximity of these journals based on Venue-Author-Coupling (producer), journal co-citation analysis (artifact), topic analysis (concept) and interlocking editorial board membership (gatekeeper) was measured. The resulting networks were examined for potential correlation using the Quadratic Assignment Procedure.

The VAC approach is used to represent the producers in this dataset. This approach measures journal proximity based on the number of authors shared by each journal pair. The VAC approach is based on the idea that an author’s choice of publication venue reflects similarity judgments authors are likely to choose venues that are thematically or socially similar.

Artifacts are measured by means of journal co-citation. This measure, introduced by McCain (1991), refers to the appearance of two journals in the same reference list of an article. The more frequently two journals appear in the same reference lists, the greater the similarity between the two journals. The journal co-citation approach measures journal proximity by the frequency with which each journal pair is co-cited by the same articles.

Topic modeling is used to capture concepts. ... The technique adopted here, the author-conference-topic (ACT) model (Tang et al., 2008), extends the LDA model by considering the author and publishing venue of the articles. LDA was developed originally as a topic modeling technique concerning the probability distribution of keywords for topics, and is particularly helpful with the ‘‘classification, novelty detection, summarization, and similarity and relevance judgment’’ of large-scale data (Blei et al. 2003, p. 993). ... This model extends the idea of LDA by taking into account the authors and publishing venues, and estimates not only the distribution of words on topics, but also the distribution of authors and venues on the topics modeled. ... Here, the outcome of the ACT model is the probability distribution of each author and each journal over topics, and the journal proximity is calculated using the cosine similarity of the journals.

The interlocking editorship approach, employed by Ni and Ding (2010), measures journal proximity based on common editorial board membership. The number of editorial board members that two journals share can be viewed as an indicator of journal similarity. ... Thus, it can be expected that if two journals have scholars in common on their editorial boards, these two journals have some degree of similarity, either cognitively or socially.

The journals were clustered using a hierarchical clustering technique with squared Euclidean distance and Ward’s method. Each journal clustering was displayed as a network (Kamada-Kawaii layout); each node (journal) was colored according to the hierarchical clustering result with the size of a
node proportional to its centrality (either degree or closeness).

Additionally, a comparison of journal proximity results was conducted using the Quadratic Assignment Procedure (QAP). QAP is commonly used in social network analysis as a means of investigating correlations between two networks. ... (Lawler 1963).

2014年6月20日 星期五

Porter, A. L., & Rafols, I. (2009). Is science becoming more interdisciplinary? Measuring and mapping six research fields over time. Scientometrics, 81(3), 719-745.

Porter, A. L., & Rafols, I. (2009). Is science becoming more interdisciplinary? Measuring and mapping six research fields over time. Scientometrics, 81(3), 719-745.

Scientometrics

本研究將跨學科研究(interdisciplinary  research)操作化的定義為:由團隊或個人從兩個或以上的知識體系(bodies of knowledge)或研究實務整合它們的觀點/概念/理論、工具/技術以及資訊/資料的一種研究模式,也就是這類研究其知識來源具有多樣性,然後分析六個研究領域在1975年和2005年的跨學科程度變化。跨學科指標的計算以引用期刊在WoS (Web of Science)上的主題分類(Subject Categories, SCs)為基礎,並且配合科學映射圖(science maps)表現科學產出在主題分類上的分散情形(dispersion)。整個分析的流程包含五個步驟:


一、將跨學科性的測量操作化。
二、建構主題分類間的相似性矩陣,做為計算整合性指標之用。
三、對相似性矩陣進行因素分析(factor analysis),將主題分類分群成為巨型學科(macro-disciplines)以便進行視覺化。
四、產生科學映射圖。
五、選取六個主題分類,做為目前的基準與未來探索。


針對操作化跨學科性的測量有幾點必須說明:首先根據Stirling的看法,探索跨學科性時,需要針對引用的學科數量、引用在學科間的分布情形、類別的相似性等面向進行研究[RAFOLS & MEYER, FORTHCOMING]。其次,本研究認為知識整合是一種認知範疇(an epistemic category),因此跨學科性指標應該建立在研究結果的內容,而不是團隊的成員,部門組織或合作上。最後,跨學科性的測量通常以引用文獻的期刊所屬的主題分類為基礎,但書目計量學的研究社群已經提出主題分類有一些問題,例如期刊叢集的研究指出僅有約50%的叢集結果和主題分類相近[BOYACK & AL., 2005; (BOYACK, personal communication, 14 September 2008)],根據引用網路得到的分類結果和主題分類之間也沒有很好的符合[LEYDESDORFF, 2006, P. 611]。但這些結果僅對科學映射圖產生有限度的影響,並且在測量整合性上,主題分類目前還是最被廣泛使用的分類資源。

本研究用來測量整合性指標[RAFOLS & MEYER, FORTHCOMING]的公式,由Rao-Stirling提出的多樣性測量方式 [STIRLING, 2007],如下:
此處pi是給定的論文上引用的參考文獻來自主題分類 i 的比例,sij是主題分類 i 和 j 的相似程度,利用cosine測量。由於許多研究 例如[GRUPP, 1990; HAMILTON & AL., 2005, or ADAMS & AL., 2007]都以Shannon或Herfindhal提出的方式測量整合性,Shannon的多樣性測量方式如
Herfindhal的多樣性測量方式如


但這兩種方法都未考慮類別間的不同;反之Rao-Stirling的多樣性則同時考慮類別數量多寡、類別上的分布平衡和類別間的相似性等三個方面。因此本研究比較此一整合性指標與Shannon和Herfindhal多樣性。

本研究以主題分類被引用的次數為資料,對每一對主題分類進行cosine測量這兩個主題分類間的相似性。當兩個主題分類被大部分的論文共同引用時,它們之間便會有很高的相似度;反之,兩個主題分類共同被引用的情形很稀少時,cosine的值接近於零。完成相似性矩陣的建立後,以主成分分析(Principal Components Analysis, PCA)進行因素分析,以最大變異量轉換(Varimax rotation)產生20個因素,將每一主題分類以其具有最高負荷的因素進行歸類。每一個因素對應一個巨型學科,某些無法歸類的主類分類另外歸於一個巨型學科,結果共有21個巨型學科。

然後以主題分類在21個因素上的負荷值為特徵,再以cosine測量主題分類之間的相似性。以Pajek將主題分類之間的相似性映射成網路圖,過濾相似性在0.6以下的連結線,做為科學映射圖。科學映射圖上呈現每一個主題分類、相對的重要性、以及彼此間的關連程度,目的在於在巨型學科間找出特定研究的主體,發現相互關連在時間上的變化以及主要的跨學科關連,更重要的是發現做為知識來源的期刊是來自於密切關連的學科或是跨越完全不同的領域。


本研究選取生物科技與應用微生物學(Biotechnology & Applied Microbiology)、電子電機工程(Engineering, Electrical & Electronic)、數學(Mathematic)、醫學(Medicine – Research & Experimental)、神經科學(Neurosciences)、物理(Physics – Atomic, Molecular & Chemical)等六個主題分類。

研究結果發現:30年間論文的平均作者數、平均參考文獻數和引用的學科數量都有很大幅度的增加,但是從跨學科指標的增加並不大。造成上述現象,可能是由於雖然引用的主題分類數量有明顯的增加,但每篇論文平均引用的參考文獻數量增加地更快,使得在不同主題上的引用比例的實際改變變得不如預期中的重要;另外,許多主題分類的引用較傾向於鄰近的主題分類,但是鄰近區域的主題分類有較高的相似值,對於多樣性的貢獻較低;最後是某些較跨學科研究的領域其測量的整合性已經到達飽和了。從科學映射圖的結果也指出論文引用的分布仍然主要集中於某些鄰近的學科領域。此外,本研究也發現Rao-Stirling的多樣性測量與Herfindhal和Shannon的測量都有很高的相關性,分別為0.91(標準差0.07)及0.88(標準差0.07) 。

Here we investigate how  the degree of interdisciplinarity has changed between 1975 and 2005 over six research domains. ... The results attest to notable changes in research practices over this 30 year period, namely major increases in number of cited disciplines and references per article (both show about 50% growth), and co-authors per article (about 75% growth). However, the new index of 
interdisciplinarity only shows a modest increase (mostly around 5% growth). Science maps hint 
that this is because the distribution of citations of an article remains mainly within neighboring 
disciplinary areas.

We measure how integrative particular research articles are  based on the association of the journals they cite to corresponding Subject Categories  (“SCs”) of the Web of Science (“WoS”)

And, we present a practical way to map  scientific outputs, again based on dispersion across SCs.

This report operationally defined interdisciplinary  research as: 
x a mode of research by teams or individuals that integrates 
x perspectives/concepts/theories and/or 
x tools/techniques and/or 
x information/data 
x from two or more bodies of knowledge or research practice. 

Our approach here is to investigate changes of degree of interdisciplinarity over time  using various established indicators (e.g. number of disciplines cited, percentage of  citations within-field), together with a new indicator developed the NAKFI evaluation  team [PORTER & AL., 2007]: 
Integration – reflecting the diversity of knowledge sources, as shown by the breadth  of references cited by a paper. 

Following Stirling’s heuristic, we have previously argued that in order to explore interdisciplinarity, one needs to investigate multiple aspects, namely: the number of disciplines cited (variety), the distribution of citations among disciplines (balance), and, crucially, how similar or dissimilar these categories are (disparity) [RAFOLS & MEYER, FORTHCOMING]. 

The computation and visualization of the interdisciplinarity measure has taken five  steps, presented consecutively in this section: 
1. Operationalization of an interdisciplinary measure (the Integration index or disciplinary diversity)
2. Construction of a similarity matrix among Subject Categories that is used to compute the Integration index
3. Grouping via factor analysis of the SCs into macro-disciplines using the similarity matrix as a base to facilitate visualization
4. Generating science maps
5. Selection of a bibliometric sample of 6 SCs, to serve as benchmarks here and in future explorations. 

In other words, since knowledge integration is an epistemic category, indicators of interdisciplinarity should be based on the content of the research outcomes rather than on team membership, departmental affiliations, or collaborations (see illustrations in case studies in RAFOLS & MEYER, 2007). 

The bibliometric community has noted that the SCs have some problems. In journal clustering exercises, only about 50% of clusters were found to be closely aligned with SCs [BOYACK & AL., 2005; (BOYACK, personal communication, 14 September 2008)]. Poor matching between SCs and classifications derived from citation networks has also been reported [LEYDESDORFF,
2006, P. 611], but surprisingly the mismatch only has limited effect on the corresponding science maps [RAFOLS & LEYDESDORFF, UNDER REVIEW].

Nonetheless, the SCs offer the most widely available categorization resource that we could ascertain for the purpose of providing an accessible measure of Integration.

As derived in RAFOLS & MEYER [forthcoming], the formula for the Integration index can be expressed as:

where pi is the proportion of references citing the SC i in a given paper. The summation is taken over the cells of the SC x SC matrix. sij is the cosine measure of similarity between SCs i and j (the cosine measure may be understood as a variation of correlation). Here this matrix sij is based on a US national co-citation sample of 30,261 papers from Web of Science as explained below in detail. 

This Integration measure (aka, Rao-Stirling’s diversity) can be compared with Shannon diversity: 

or with Herfindhal’s diversity (the complement of Herfindahl’s concentration):

The power of the Integration index is that it characterizes interdisciplinarity in terms of the diversity of knowledge sources of papers, using a general formulation of diversity [STIRLING, 2007] rather than an ad hoc indicator.

A number of researchers have used these traditional measures of diversity, such as Shannon or Herfindhal, to measure interdisciplinarity [E.G. GRUPP, 1990; HAMILTON & AL., 2005, or ADAMS & AL., 2007]. These measures do not take into account how different the categories are, whereas our Integration measure reveals increased diversity only when added categories are significantly different.

In particular, a broad national sample of articles from WoS is used to create the sij matrix that underlies the metrics used for computing Integration. First we describe the sample used as a basis for the similarity matrix; second, the construction of the matrix.

We combine six separate weeks of all papers in WoS, with one or more authors having a USA address, sampled during 2005–2007, to obtain 30261 articles. This provides a broadly based, yet manageable base sample. We processed the “Cited References” of these abstract records to identify the “Cited SCs.”

Our sample of 30261 WoS articles contains 1,020,528 cited references (an average of 33.7 per article). Of those, our thesauri link 768,440 to a particular Subject Category. Another 28,000 have been checked and assigned to “not being in an SC.”

For our purposes in addressing cited SCs, the list includes a few more than the current set, for a total of 244 SCs. The sample contains 1,114,930 instances of cited SCs.

The 30261 articles, by 244 SCs, described allow for construction of a co-citation similarity matrix, sij, using Salton cosine [SALTON & MCGILL, 1983; AHLGREN & AL., 2003].

The values of sij are high (i.e. closer to one) when SCs i and j are co-cited by a high proportion of articles that cite one or the other. The cosine value approaches zero when two SCs are rarely cited together.

For various purposes and in particular for visualization, it helps to consolidate the narrow research areas of the ISI SCs into larger categories, which we call “macro-disciplines.”

We base our grouping of SCs on a type of factor analysis – Principal Components Analysis (PCA) – following a similar methodology to that developed by LEYDESDORFF & RAFOLS [2008] to cluster SCs into macro-disciplines.

Within VantagePoint, we constructed the matrix of cosine similarities for the 244 cited SCs by 244 cited SCs described in the previous section. ... We explored various factor analysis solutions, eventually adopting a 20-factor solution (Varimax rotation). ... The 21 macro-disciplines reflect this factor solution.

So, to a considerable degree, named sub-disciplines do not fully coalesce within a single macro-discipline. This warns that the evolving research enterprise does not neatly conform to the traditional scholarly disciplines.

These maps present the SCs, their relative importance in size, and how related they are to each other over all science. The main aim of these science maps is to locate particular bodies of research among the macro-disciplines. ... That can help identify changes in degree of interrelationship over time, and key cross-“disciplinary” relationships that might benefit from nurturing. It should also be informative to see whether knowledge sources of a set of publications are coming from research domains that are closely related (little interdisciplinarity) or that span very disparate domains (high interdisciplinarity). 

We then construct a new Salton cosine similarity matrix among SCs using the loadings of each SC on the 21 factors (as discussed in the previous subsection). This matrix is then uploaded into the network analysis software Pajek [BATAGELJ & MVAR, 2008]. In Pajek, the minimum similarity threshold was arbitrarily set to 0.6 (this choice was found to provide a good readability-to-accuracy trade-off) and the SCs were distributed in a 2-D plane according to their similarities, to obtain a base science map.

Since research collaboration is often (and sometimes mistakenly) associated with interdisciplinarity, we examine measures of co-authorship. ... However, within research domain, the number of authors per paper has escalated remarkably, with about 75% average growth. This increase ranges from 48% in Math and 54% in Physics-AMC to 90% in Neurosciences. 

Before turning to Integration scores, we consider the number of distinct SCs that one article cites. ... Table 2 and Figure 4 show a sturdy increase in the breadth of citing in all six of these research domains (about a 50% growth on average). 

Integration scores are tabulated in Table 2 and shown in Figure 5. We see that over time, there is a modest increase in Integration scores and that math researchers are notably less integrative in their citing patterns. However, math has the highest relative growth (39%) whereas other SCs’ growth ranges from 3% to 14% (5% on average). t-tests between the 1975 and 2005 samples show these differences to be highly significant (<.005 for EE, assuming either equal or unequal variances; all others even more highly significant).

Pearson’s correlation between Integration and Herfindhal takes a mean value of 0.91 (standard deviation = 0.07) and between Integration and Shannon, a mean value of 0.88 (standard deviation = 0.07). These high correlations confirm that Integration is very closely associated with traditional diversity indicators – as could be expected by construction.

The main finding is that Integration scores increase over time, but significantly less so than other indicators, such as percentage of single-authored papers, mean authors per paper, and mean number of disciplines per paper.

First, although the number of cited SCs increases significantly, since the average number of references in a paper also shows a quicker increase (see central columns in Table 2), the actual change in the proportions of citation to different SCs is not as important as could be expected.

Second, as we will show in Figures 7 through 10, the citation patterns of a given SC tend to be with SCs in its vicinity. Since these neighboring SCs have high similarity values with the one investigated, their contribution to Integration (to diversity) is smaller than in other indicators. This means that the Integration score “deflates” the diversity recorded by Shannon or Herfindahl because most of the cited SCs are not very different from the SC doing the citing.

This is much easier to convey using science maps that directly show the three aspects of disciplinary diversity, namely:
1. the variety of “disciplines” (i.e., discrete research areas, the SCs, shown by the number of nodes in the map)
2. the balance, or distribution, of disciplines (relative size of nodes)
3. the disparity, or degree of difference, between the disciplines (distance between the nodes)

These maps were created followed the techniques developed in LEYDESDORFF & RAFOLS [2008], in the context of the current interest in science mapping [MOYA-ANEGON & AL., 2004; BOYACK & AL., 2005; MOYA-ANEGON & AL., 2007]. ... In the figures presented in this article, we only label groups of SCs on the basis of macro-disciplines found by factor analysis, as explained in the methodology. 

However, the perspective provided by the Integration score and the science maps suggests that the practice of interdisciplinarity in citations occurs mainly between neighboring SCs and has undergone a much more modest increase (on average only 5%, excluding math).

This is mainly for two reasons: first, although the number of cited SCs has increased, the growth of citations means that the increase in the proportion of citations to new SCs is small; second, the newly cited SCs tend to be in the vicinity of the previous ones – hence they don’t add as much interdisciplinarity as they would if they were very disparate/distant disciplines. Moreover, for already very interdisciplinary SCs, such as Neuroscience, the indicator may have a certain “saturation” effect. 



2014年1月24日 星期五

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

本研究探討在地理上的生產力遷移(shifts in productivity)是否發生在書目計量學(bibliometrics)、資訊計量學(informetrics)和科學計量學(scientometrics)等計量學(metrics)領域,也就是歐洲的貢獻明顯地成長,並且北美的貢獻相對來說有減少的情形。

有關計量學的研究,Hood and Wilson (2001)和Stock and Weber(2006)等研究都分析了這個領域的文獻成長情形。Hood and Wilson (2001)回顧了計量學領域的發展,並且比較bibliometrics、scientometrics和informetrics的相關文獻,發現bibliometrics還是在相關領域上使用最廣泛的詞語。Stock and Weber(2006)從觀察中確認這個領域從1980年後便持續地成長。Wolfram (2008)則發現在計量學領域中,北美的文獻有明顯地減少而歐洲則是急遽地增加的情形。

本研究利用bibliometrics、scientometrics、informetrics、cybermetrics、webometrics、citation analysis、link analysis和citation indexes做為檢索的問句,同時再加上Scientometrics和Journal of Informetrics兩種期刊的論文,從Web of Science資料庫中進行檢索。結果共檢索出4404筆論文資料。

在這些論文資料裡,共有75個國家。以地區來區分,歐洲在每個時段上具有最大的貢獻,不論是數量或所占比率都有成長,亞洲所佔的相對比例在22年間有很大的成長,北美雖然在數量上有成長,可是相對的比例呈現緩慢的下降。每個地區的作者會偏好在本身地區的期刊上發表,舉例而言,歐洲作者發表論文的前五個期刊中有四個歐洲期刊,南美也有類似的情形,但是亞洲的情形例外,前五個期刊中有四個是歐洲期刊,另一個則是北美的期刊。

自1990年代中期後,國家間的合作情形增加許多,之前國際合作的論文每年為1到19篇,2008年已大幅增加為96篇。美國是國際合作佔最多的國家,但以地區來說,歐洲平均每個國家的國際合作數為5.78篇論文,多於世界其他部分的4.47篇論文。

此外,歐洲則有許多具有國際合作經驗的機構,共有16所研究機構有國際合作經驗,北美則有8所,亞洲有1所。機構間的合作來說,在1987年每篇論文平均只有1.1個機構,但在2007年則增加為1.96。

本研究且利用MDS、VOSviewer和Pajek將這些論文上的國家與機構之間的合作關係,呈現為圖形。

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).


This investigation was prompted by interest in whether shifts in productivity based on geography are observed in the bibliometrics, informetrics and scientometrics areas.

One of the authors conducted a pilot study to determine whether there have been clear declines in North American contributions to the metrics literature base (Wolfram, 2008). The author found that there was indeed a notable relative decline in North American contributions and a sharp increase in European contributions.

Hood and Wilson (2001) examined the growth of literature of the metrics area. They provided an historical treatment of the development of these areas that included earlier studies of the field. In their research, literature associated with bibliometrics, informetrics and scientometrics was compared for the period 1968–2000. The authors noted that bibliometrics was still the most widely used term for metrics research.

More recently, Stock and Weber(2006) conducted a Web of Science search for records specifically including metrics terms and allied areas. They observed contributions had grown substantially since 1980.

Search parameters included the Boolean ORed result of bibliometrics, scientometrics, informetrics, cybermetrics and webometrics, in truncated form (e.g., webometri*), along with the phrases “citation analysis”, “link analysis” and “citation indexes”. ... These search results were ORed with the two primary journals that publish metrics research that are indexed by WoS, namely Scientometrics and the Journal of Informetrics.

A pair-wise comparison of all collaborations at the national and institutional levels was then conducted from which a cooccurrence matrix could be compiled.

Multidimensional scaling (MDS) analysis was used to visualize the relationships among countries. Because the data represent a type of similarity measure represented as a symmetric matrix, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006).

The recently developed visualization tool VOSviewer (van Eck &Waltman, 2010) was also used to provide an alternate visualization of the relationship outcomes. Like MDS, VOSviewer (http://www.vosviewer.com/) relies on a distance-based approach to mapping informetric relationships. Instead of using more traditional similarity measures to produce a normalized outcome for co-occurrences as used in MDS, relationships are based on association strengths, so the algorithm is somewhat different than PROXSCAL and, therefore, can produce different outcomes. Details of the comparison of different measures can be found in van Eck and Waltman (2009).

The network visualization software Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/) was used as well. Unlike the distance-based mapping of PROXSCAL and VOSviewer, Pajek produces directed or undirected network maps, with the strength of the relationships represented by the thickness of connecting lines between vertices on the map. Distances are used more for clarification, but proximities do not necessarily indicate a stronger relationship.

The search parameters retrieved 4404 publications.

Europe shows the highest levels of contribution, both in absolute and relative terms over the time period of the study. Growth patterns in absolute terms are nonlinear based on trend line analysis in MS Excel; however, the R-squared goodness-of-fit values for even the best fitting models (higher order polynomials) were never more than 0.95, indicating a less than desirable fit.

Relative contributions based on geographic divisions have been largely stable. An exception is Asia, which had an increasing relative contribution over the 22-year time frame of the study. Although North American contributions have continued to increase in absolute numbers, the relative contribution shows a slow average decline over time.

The top five journals listed for each continent demonstrated a regional preference for publication outlets from that region. So, for example, four of the top five journals for European publications were published in Europe, and four of the top five journal outlets for South America were South American. The exception to this was Asia. Four of the top five journals for Asian publications were European and one was North American. This outcome may be a reflection of the data extraction method, the indexing practices of WoS, or a preference during the study time frame for Asian scholars to publish in Western journals.

Seventy-five countries were represented in the record set.

The number of metrics papers published annually that represent collaborations between two or more countries has increased greatly since the mid-1990s. Prior to this time, the number of internationally collaborative papers ranged from 1 to 19 papers annually. Over the last decade this number has increased to a high of 96 papers in 2008.

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).

Sixteen of the institutions on the list are European, eight are North American, and one is Asian. The United States has the largest number of institutions represented (five), followed by Belgium (four – note: one institution merged with another institution to form a new entity).

There has been steady growth in inter-institutional collaboration over the 22 years. The mean number of collaborative institutional partners within the dataset has steadily increased from a low mean of 1.1 institutions per publication in 1987 to a high of 1.96 institutions per publication in 2007.

Europe, and in particular Western Europe, clearly dominates in the production of metrics literature. The United States continues to be the largest singular contributor, but this appears to be changing. North American contributions as a whole continue to increase, but represent a smaller percentage of worldwide production. European contributions have grown tremendously, especially during the last 5 years of the study period. This same period is marked by impressive growth from Asia.

It should be noted that WoS increased its coverage in 2008 by including more regional journals. These inclusions possibly could contribute to the increase in Asian contributions, but the observed growth for Asia was already evident prior to any such additions.

International and inter-institutional collaborations do not necessarily reveal strong geographic affinities, although the multiple institutional affiliations by a number of scholars associated with Flemish institutions do contribute to the strengthening of regional ties. Undoubtedly, the growth of the Internet and increasing availability of other telecommunication technologies have made these collaborations less distance dependent.

2014年1月18日 星期六

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

本研究利用社會網絡分析、共現分析(co-occurrence analysis)、叢集分析和詞語的頻率分析等多種分析技術,從Scientometrics期刊1978到2004年發表的1927筆論文資料,探討科學家合作網絡的結構特性、整個網絡上的合作領域以及個別的合作網絡、合作網絡上的合作中心(collaborative  center)。

過去的研究裡,Schubert (2002) 和 Dutt, Garg, & Bali (2003)都是針對國家間合作的巨觀層次。Kretschmer (2004) 認為巨觀和中觀(meso)層次的分析無法足夠地反映個人之間的合作趨勢,因此呼籲應在微觀層次的分析投注更多努力。

1927筆論文資料裡,單一作者的論文共有1052筆,所以仍稍占多數。作者數大於3的論文僅占非單一作者論文的13.71% (120/875),顯然研究Scientometrics的團隊規模都不大。發表3篇論文以及以上的高生產作者共計234人,其中有69.66%的作者曾發表與其他作者合作的論文。將這些作者間的合作關係表現成網絡,並利用Bibexcel對這個網絡上的節點進行叢集分析,共發現22個叢集。前兩個較大的叢集分別有15與14個科學家。網絡上最大的相連成分上共有15個叢集,共有合作經驗的高生產作者中的96位,占58.90%。合作網絡共有401條連結線,網絡密度為0.03,顯示Scientometrics領域的合作很鬆散。

對每一個節點計算它們的三種中心性,結果發現中心性和對應作者的生產力之間有很顯著的正相關,表示高生產力的作者同時也活躍在Scientometrics領域的合作網絡上。其中Glänzel的程度中心性最高,總共和其他18位作者有合作關係。

以詞語的頻率分析每個叢集的主題,最大的兩個叢集有類似的主題,但使用的研究方法略有不同。此外,研究主題為科學合作的四個叢集間幾乎沒有連結,同樣的情形也發生在研究科學與技術之間關係的四個叢集。

The structure of scientific collaboration networks in scientometrics is investigated at the level of individuals by using bibliographic data of all papers published in the international journal Scientometrics retrieved from the Science Citation Index (SCI) of the years 1978–2004.

Combined analysis of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of words is explored to reveal: (1) The microstructure of the collaboration network on scientists’ aspects of scientometrics; (2) The major collaborative fields of the whole network and of different collaborative sub-networks; (3) The collaborative center of the collaboration network in scientometrics.

Schubert [8] and Dutt etc. [9] presented international collaboration characteristics in the scientometrics community itself, focusing on country aspects at macro level.

Kretschmer [6] appealed to devote more efforts to investigations at micro level in the future because the knowledge at meso and macro level does not yet adequately reflect the trends in cooperation between individuals.

The study is based on bibliographic data retrieved from the Web of Science. The data contains all types of documents published in Scientometrics during 1978 to 2004.

In this study we have adapted an integrated procedure of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of title words.

Bibexcel is designed as a tool for manipulating bibliographic data, which is a free online-software published by Persson. In the present study, Bibexcel is used to do cooccurrence analysis and cluster analysis.

Following the methods of Otte & Rousseau [11], White [13] and Kretschmer & Aguillo [12], SNA was applied to display the microstructure of collaboration networks in scientometrics with Pajek.

Moreover, we used frequency analysis of title words to display the main collaborative field of different sub-networks. The software for frequency analysis is demo version of Wordsmith Tools published by Oxford University Press and available online.

There were 1927 documents published in Scientometrics during 1978 to 2004 (see Table 1).



From Table 1, we found that the pattern of co-authorship was still dominated by single-authored papers as the conclusion drawn by Dutt etc. [9].

While the number of multi-authored papers (the number of co-authors is more than 3) accounts for 13.71% only, which indicates that team size in scientometrics is not large.

In order to show the main structure of the network, each author must published 3 papers or more to be included in this integrated analysis. This threshold resulted in a total of 234 prolific authors publishing 3 or more papers during 1978 to 2004, among them there are 163 authors published co-authorship papers, accounting for 69.66% of the prolific authors.



Based on cluster analysis embedded in Bibexcel, we gained 22 clusters circled by solid lines (see Figure 1). We identified these clusters as sub-networks in the field of scientometrics.

The largest subnetwork is number 1 that has 15 collaborators, and the second largest one is number 2, which has 14 collaborators, and so on.

We noticed that there was totally 15 subnetworks connected with each other composing the largest central component, which had 96 numbers accounting for 58.90% of the prolific authors published co-authorship papers.

Density is an indicator for the general level of connectedness of the graph. ... In the present study, there are totally 401 links in the network, so the density of the network is 0.03, which indicates that the collaborative network in the field of scientometrics is very loose.

So an author who has high degree centrality must has collaborated with many other authors, which means the author is a central collaborator of the whole network. In the present study, Glänzel who has 18 co-workers is the central author of the whole network.

We found a positive and significant correlation between output of authors and the centrality measures (r=0.648, 0.437, 0.338 respectively at the 0.01 level, see Table 4) after investigating the correlations between output and the three centralities of the 125 authors in the 22 sub-networks, which indicated that most of the prolific authors are also active in collaboration network in the field of scientometrics.

We have also presented the main collaborative field of different sub-networks in scientometrics and found that the two biggest sub-networks have the similar collaborative topic with slightly methodological difference. In addition, we found an interesting phenomenon that four sub-networks dealing with scientific collaboration didn't collaborate with each other except sub-network 3 and 12. Moreover, four subnetworks studying technology and science never collaborated with each other at all.

2013年4月21日 星期日

Leydesdorff, L., & Rafols, I. (2009). A global map of science based on the ISI subject categories. Journal of the American Society for Information Science and Technology, 60(2), 348-362.

Leydesdorff, L., & Rafols, I. (2009). A global map of science based on the ISI subject categories. Journal of the American Society for Information Science and Technology, 60(2), 348-362.

information visualization

本研究對175個ISI主題分類(ISI subject categories)的相互引用資料進行探索性因素分析(exploratory factor analysis),從這個結果來驗證ISI主題分類是否能夠進行資訊視覺化。相關的研究包括Boyack, Klavans, and Börner (2005)利用VxOrd的VxInsight演算法將所有ISI資料庫收錄的期刊,根據它們相互的引用關係進行視覺化,能夠反映科學結構的期刊映射圖,同時以k-means叢集演算法產生新的分類;Moya-Anagón et al. (2007)則是利用ISI主題分類的共被引(co-citation)資料做為尋徑者網路(Pathfinder network)演算法的輸入資料來產生映射圖。本研究的研究資料是2006年的ISI期刊引用資料,在2006年175個ISI主題分類內有三個主題分類(“Psychology, biological,” “Psychology, experimental,” and “Transportation”)沒有引用資料,但有被引用的資料。資料先以cosine進行正規化,利用SPSS (v15)進行因素分析,並且利用Pajek程式 (Batagelj & Mrvar, 2007)的Kamada and Kawai (1989)演算法進行視覺化。以具有引用資料的172個主題分類而言,經過因素分析後,共可發現14個因素,這些因素與學科(disciplines)的分類相符合。以被引用的資料進行因素分析,其結果也可以看出學科分類的樣貌。比較引用資料和被引用資料的結果,172個主題分類的154個(89%)在兩個結果中落於相同的因素。

The ISI subject categories classify journals included in the Science Citation Index (SCI). The aggregated journal-journal citation matrix contained in the Journal Citation Reports can be aggregated on the basis of these categories. This leads to an asymmetrical matrix (citing versus cited) that is much more densely populated than the underlying matrix at the journal level. Exploratory factor analysis of the matrix of subject categories suggests a 14-factor solution. This solution could be interpreted as the disciplinary structure of science. The nested maps of science (corresponding to 14 factors, 172 categories, and 6,164 journals) are online at http://www.leydesdorff.net/map06.

In contrast, Boyack, Klavans, and Börner (2005) used the VxInsight algorithm (Davidson, Hendrickson, Johnson, Meyers, & Wylie, 1998) in order to map the whole journal structure as a representation of the structure of science.

Moya-Anagón et al. (2004, 2007) used cocitation and PathFinder for mapping the whole of science on the basis of the ISI subject categories.

Klavans and Boyack (2007, p. 438) noted that a journal may occupy a different position in a different context: Many journals report on developments in multiple disciplines; journals can also function as a major source of references in more than one specialty.

Since citation relations among journals are dense in discipline-specific clusters and otherwise virtually nonexistent, the journal-journal citation matrix can be considered nearly decomposable.

The next-order units represented by the square submatrices— and representing in this case disciplines or specialties—are reproduced in relatively stable sets (of journals), which may change over time. The sets of journals are functional subsystems that show a high density in terms of relations within the center (i.e., core journals), but are more open to change in relations at the margins.

The decomposition into nearly decomposable matrices has no analytical solution. However, algorithms can provide heuristic decompositions when there is no single unique correct answer (Newman, 2006a, 2006b).

The number of category attributions in the Science Citation Index is 9,848 for 6,164 journals in 2006 or, in other words, approximately 1.6 categories per journal. The coverage of the 172 categories ranges from 262 journals sorted under “Biochemistry and Molecular Biology” to 5 journals sorted under a single category. The average number of journals per category is 56.3 (see Figure 1).

In other words, our research question is different from Boyack et al.’s (2005) effort to generate a new classification using a bottom-up strategy and from that of Moya-Anagón et al. (2007), who employed the ISI subject categories as units of measurement (at p. 2169), and used factor analysis of the cocitation matrix for the validation of their so-called “factor scientograms.”

We wish to question the quality and validity of using the ISI subject categories for mapping purposes. Can these subject classifications be used in further research to demarcate the sciences and perhaps as field delineations, and if so, under what conditions?

As noted above, Moya-Anegón et al. (2007, p. 2173) used factor analysis of the cocitation matrix of the 218 categories of the Science Citation Index and Social Science Citation Index (2002) combined for the validation of their visualizations. These authors stated that a scree test had led them to the choice of 16 factors.

We approach the problem first factor-analytically using the asymmetrical matrix of aggregated citations among categories, and will subsequently try to map the sciences hierarchically top-down insofar as our results show that it is legitimate for us to do so.

The data was harvested from the CD-ROM version of the Journal Citation Reports of the Science Citation Index 2006. As indicated above, 175 subject categories are used. Three categories (“Psychology, biological,” “Psychology, experimental,” and “Transportation”) are no longer used as classifiers in the citing dimension, but four journals are still indicated with these three categories in the cited dimension. Thus, we work with 172 citing and 175 cited categories.


The matrix, accordingly, contains two structures: a cited and a citing one. Salton’s cosine was used for normalization in both the cited and citing directions (Ahlgren, Jarneving, & Rousseau, 2003; Salton & McGill, 1983).


Pajek is used for the visualizations (Batagelj & Mrvar, 2007) and SPSS (v15) for the factor analysis. The threshold for the visualizations is pragmatically set at cosine≥0.2. Visualizations are based on the algorithm of Kamada and Kawai (1989).



Let us focus on the structure in the citing dimension because this structure is actively maintained by the indexing service and is therefore current.... The factor loadings for the 172 categories on the 14 factors in the citing dimension are provided in the Appendix. They can be interpreted in terms of disciplines, such as physics, chemistry, clinical medicine, neurosciences, engineering, and ecology.

The factors in the cited dimension can be designated using precisely the same disciplinary classifications, but their rank order (that is, the percentage of variance explained by each factor) is different (Table 2). Out of the 172 categories, 154 (89%) fall in the same factor in both the citing and cited projections. ... The strong overlap between the results of the factor analysis in the cited and the citing dimension (Table 2) suggests that the matrix is nearly decomposable in terms of central tendencies.

Our results are consistent with previously reported maps (Boyack et al. 2005; Boyack & Klavans, 2007; Moya-Anagón et al., 2007), but we chose to exclude the social sciences.