顯示具有 VOS 標籤的文章。 顯示所有文章
顯示具有 VOS 標籤的文章。 顯示所有文章

2017年11月30日 星期四

Chen, S., Arsenault, C., Gingras, Y., & Larivière, V. (2015). Exploring the interdisciplinary evolution of a discipline: the case of Biochemistry and Molecular Biology. Scientometrics, 102(2), 1307-1323.

Chen, S., Arsenault, C., Gingras, Y., & Larivière, V. (2015). Exploring the interdisciplinary evolution of a discipline: the case of Biochemistry and Molecular Biology. Scientometrics102(2), 1307-1323.

本研究從跨學科性(interdisciplinarity)的變化、核心學科的確定、學科的出現和潛在的學科檢測幾個方面,探討生物化學與分子生物學(Biochemistry and Molecular Biology, BMB)一百年間的科際整合演化,並利用科學地圖和河流圖(StreamGraph)做為檢視科際整合演化的視覺化工具

Porter和Rafols(2009)研究六個研究領域的三十年間跨學科性的演變,研究結果顯示,在這段時間內有顯著的變化,特別是每篇文章引用的學科數量和參考文獻的數量,以及每篇文章的合著者數量。不過,Rao-Stirling指數只是略有上升,Porter和Rafols認為這是由於論文的引用還是主要是在相鄰的學科領域內。Larivie`re and Gingras (2014)的研究發現,在1945和1975間跨學科研究下降後,已經逐漸上升。Levitt et al. (2011)研究SSCI (Social Science Citation Index)的各主題類別的跨學科性,則是發現1980和1990之間下降後,2000年已恢復1980年的水準。

在文獻計量學,跨學科性的測量通常是根據跨學科的引文關係。例如,Porter和Chubin(1985)使用「外部類別引用」(Citations Outside Category, COC)的指標衡量文章的跨學科性程度,這個指標定義為來自不同學科的引用或參考文獻的百分比。這種方法被應用到多個研究,例如Rinia等人(2002a)測量所有科學領域的跨學科性,Morillo et al. (2001) 測量化學領域的跨學科性。Adams等人(2007)除了不同學科的引用參考文獻的比例,也使用了引用的學科數量和Shannon多樣性指數 (Shannon Diversity Index)。Carley and Porter (2012)利用Rao-Stirling多樣性指數(Rao-Stirling diversity index),透過引用文獻,探索知識整合。Levitt和Thelwall(2008)則是使用WoS和Scopus的主題分類,以論文發表期刊的多重分類測量跨學科性

Rinia等人(2002b,244)將跨學科性定義為一組研究人員在他們的“主要”學科之外發表的文章的百分比。Qiu (1992)使用作者的組織隸屬關係探究跨學科合作。Abramo等人(2012)以義大利為案例研究,以研究人員的學科作為類別指定的方法,利用不同學科的合作者之間的合作確認跨學科性。Le Pair(1980)和Sugimoto等(2011)則是利用作者最高學歷學科探索跨學科Le Pair(1980)將跨學科合視為科學家在他們的生涯中從一門學科向另一門學科的轉移。Sugimoto等(2011)利用80年(1930-2009)的博士論文,從學術譜系描述了圖書情報學跨學科性變化程度。

本研究利用WoS資料庫上主題類別BMB下從1910年起所有的文章,共 1,539,526篇,以及其文件類型為期刊論文的參考文獻,共40,855,852 筆。參考文獻的期刊以美國國家科學基金會(National Science Foundation, NSF)的主題分類作為學科指定的系統,共143個分類,每種期刊主要被指定到一個分類,但研究上僅考慮超過250筆參考文獻的類別,並且以參考文獻所占比例較多的類別做為核心學科,此外也預測有潛力的學科。計算上使用兼顧引用文獻分布的學科數量和集中性(concentration)以及學科間相似性Rao-Stirling指數 (Rafols and Meyer, 2010)來分析COC。本研究使用相對開放性(relative openness)指標(Lee等人, 2009)來衡量引用學科的影響,並通過查看這些學科的影響的變化來探索科際整合的演變。

在視覺化部分,本研究使用河流圖表現核心學科長期變化,河流的寬度變化表示學科影響力的增加或減少。將按照時間順序出現學科可視化是另一種說明BMB科際整合的方法,本研究選擇引用次數等於或超過250次的學科探究BMB的新興學科。另外,本研究以VOSveiwer產生科學地圖,圖上的每一個單位為NSF主題分類的143個學科,其距離與大小由10年間(2003~2012年)的學科共被引矩陣計算,並依據NSF的14個主要學科分類著上顏色。




本研究的結果顯示,在100年間每年BMB引用的學科從1成長到93,跨越12個主要學科分類,只有Humanities和Arts兩個主要學科分類未曾被引用主要學科分類中較重要的有Clinical Medicine、Biomedical Research和Biology,另外,Physics和Chemistry也相當重要。從引用學科增長的數量,可見BMB在100年期間愈來越具有跨學科性。


比較各學科被BMB引用的比例(ri)以及它們的論文占所有發表論文的比例(pi),本研究確認出15個核心學科:



核心學科的首次引用大約都出現在1950年之前,較晚出現的核心學科大都是新興的學科。並且在科學地圖上,核心學科通常被映射在BMB的附近,也就是說與BMB較近的學科在早期便被BMB引用。

雖然BMB的引用主要來自本身(45.32%),但引用的比例明顯在下降中,1910年為74.0%,但2012年只有32.1%。除了BMB本身以外,一般生物醫學研究(General Biomedical Research)是另一個最多引用的學科。然而隨時間改變,不同的時期有不同重要的科際整合學科。在早期,生理學(Physiology)、藥學(Pharmacology)和免疫學(Immunology)是最重要的學科,然而近期這些學科的重要性減少,取而代之的是細胞生物學(Cellular Biology)、細胞學和組織學(Cytology and Histology)、遺傳學(Genetics and Heredity)與癌症學(Cancer)等多種學科,以及一般生物醫學研究。










在剛開始的時期,BMB主要引用本身學科的論文,然後是化學(Chemistry, 1924)、臨床醫學(Clinical Medicine,
1937)和生物學(Biology)。這三個學科也是在整個NSF分類系統所形成的科學地圖上與BMB最接近的學科。隨著BMB的發展,距離較遠的學科也逐漸加入。根據BMB的跨學科發展順序,可以分為四個時期:第一時期為1910到1960年的50年間,這個時期共有來自生物醫學、臨床醫學、化學和生物學的18個學科,在這個時期內核心學科大多已經加入。第二時期為1961~1981年的20年,這時期新增了34個學科,這時期除了生物醫學和臨床醫學的引用大量增加之外,特別值得一提的是來自物理學的參考文獻增加,顯示BMB的跨學科範圍開始從它鄰近的學科向外擴大。第三時期則是從1982到2002年的20年,這個時期新增了27個學科,跨學科的範圍更擴大到工程與科技(Engineering and Technology)、心理學、地球與空間和數學。第四時期則是從2003年開始,共有16個新學科,特別一提的是這個時期加入的圖書資訊學(Library and Information Science),其原因是這個時期開始有研究利用書目計量學方法對BMB的研究進行評估。從四個時期的科學地圖也可以發現,BMB引用的學科愈來愈多元,而且從一開始使用的學科都是在科學地圖上較接近的學科,愈到後期愈多來自距離較遠的學科。




以下則是最近十年增長最快的前五學科,這些學科可以視為是具有與BMB具有跨學科潛力的學科。




在100年間,Rao-Stirling指標從約0.3變為約0.6,成長約2倍,與Porter and Rafols (2009)的研究相比,同一時間區間(1975~2005年),Porter and Rafols (2009)研究的6個學科,Rao-Stirling指標並沒有明顯增加,但BMB則約增加0.32倍。

本研究證實跨學科主要從相鄰領域向認知較遠的領域演變,並且BMB研究者引用其他學科文獻日益增加,而雖然引用較遠領域的比例較小,但正在顯著增加,因此是BMB較有跨學科潛力的學科



2014年9月10日 星期三

Waltman, L., Yan, E., & van Eck, N. J. (2011). A recursive field-normalized bibliometric performance indicator: An application to the field of library and information science. Scientometrics, 89(1), 301-314.

Waltman, L., Yan, E., & van Eck, N. J. (2011). A recursive field-normalized bibliometric performance indicator: An application to the field of library and information science. Scientometrics, 89(1), 301-314.

在利用引用做為研究成效測量指標的基礎時,一般有兩種方法,一種是根據領域分類系統(field classification scheme)使引用次數正規化,另一為遞迴式的引用加權(recursive citation weighting)。前者認為不同領域有不同的引用密度(每一出版品上引用的平均數目),因此需要進行調整。在這種方法裡,調整的方式有兩種:一種是根據出版品所指定的領域進行調整,另一種則是從出版品引用的參考文獻數進行調整,這種方式又稱為來源正規化(source normalization)。後者的遞迴式指標則是認為從有影響力的出版品、有聲譽的期刊以及知名作者處來的引用較為有價值。

過去的研究,根據它們是領域分類系統或是來源正規化已是是否為遞迴式的加權機制分為下面的四種:


可以發現目前並沒有根據分類系統正規化同時採用遞迴式機制的方法,因此本研究便式提出一個根據分類系統正規化同時採用遞迴式機制的引用評估方法。

本研究以圖書資訊學為範例,分析這個領域的期刊以及研究機構的引用影響力。在這裡,圖書資訊學領域是以Journal of the American Society for Information Science and Technology (JASIST) 為種子期刊,以共同引用為基礎,搜尋和JASIST有最強關連的期刊,同時這些期刊也必須被歸類在並且在WoS的資訊科學和圖書館學(information science & library science)主題分類下,結果共有47種期刊。連同JASIST在內的48種期刊便做為圖書資訊學領域的代表。並且針對這些期刊利用它們之間的書目耦合(bibliographic coupling)資料以及VOS叢集演算法歸類成三群,包括圖書館學(Library Science)、資訊科學(Information Science)以及科學計量學(Scientometrics),期刊的題名與它們的歸類情形,如下表所示:


選取圖書資訊學領域48種期刊從2000到2009年發表的類型為Article或Review的文章進行分析,共12202篇。


首先將圖書資訊學全體視為是一個領域。TABLE 4是根據MNCS評估指標的前十種圖書資訊學期刊,將α分別設為1(也就是沒有遞迴式的機制)及20(遞迴的結果收斂)。在α=1的情形,前十種期刊主要為資訊科學及資訊計量學,圖書館學僅占有三種(4, 8與10)。在α=20時,前十種期刊則幾乎是資訊科學及科學計量學,圖書館學僅有排名第9的一種。同樣以MNCS評估指標來評估研究機構時,也可以發現以科學計量學為主要研究項目的機構在α=20時,獲得更好的排名。


當將圖書資訊學分為三個次領域計算時,從TABLE 6的前十種圖書資訊學期刊,可以看到三個次領域的分布已經比較平衡。


scientometrics

Two commonly used ideas in the development of citation-based research performance indicators are the idea of normalizing citation counts based on a field classification scheme and the idea of recursive citation weighing (like in PageRank-inspired indicators).

Our empirical analysis shows that the proposed indicator is highly sensitive to the field classification scheme that is used. The indicator also has a strong tendency to reinforce biases caused by the classification scheme.

One stream of research focuses on the development of indicators that aim to correct for the fact that the density of citations (i.e., the average number of citations per publication) differs among fields.

One approach is to normalize citation counts for field differences based on a classification scheme that assigns publications to fields (e.g., Braun and Gla¨nzel 1990; Moed et al. 1995; Waltman et al. 2011). The other approach is to normalize citation counts based on the number of references in citing publications or citing journals (e.g., Moed 2010; Zitt and Small 2008). The latter approach, which is sometimes referred to as source normalization (Moed 2010), does not need a field classification scheme.

A second stream of research focuses on the development of recursive indicators, typically inspired by the well-known PageRank algorithm (Brin and Page 1998). ...  The underlying idea is that a citation from an influential publication, a prestigious journal, or a renowned author should be regarded as more valuable than a citation from an insignificant publication, an obscure journal, or an unknown author.

It is sometimes argued that non-recursive indicators measure popularity while recursive indicators measure prestige (e.g., Bollen et al. 2006; Yan and Ding 2010).

To test our recursive MNCS indicator, we use the indicator to study the citation impact of journals and research institutes in the field of library and information science (LIS).

We focus on the period from 2000 to 2009. Our analysis is based on data from the Web of Science database.

We first needed to delineate the LIS field. We used the Journal of the American Society for Information Science and Technology (JASIST) as the ‘seed’ journal for our delineation. We decided to select the 47 journals that, based on co-citation data, are most strongly related with JASIST. Only journals in the Web of Science subject category Information Science & Library Science were considered. JASIST together with the 47 selected journals constituted our delineation of the LIS field.

From the journals within our delineation, we selected all 12,202 publications in the period 2000–2009 that are of the document type ‘article’ or ‘review’.

We first collected bibliographic coupling data for the 48 journals in our analysis. Based on the bibliographic coupling data, we created a clustering of the journals. The VOS clustering technique
(Waltman et al. 2010), available in the VOSviewer software (Van Eck and Waltman 2010),
was used for this purpose. We tried out different numbers of clusters. We found that a solution with three clusters yielded the most satisfactory interpretation in terms of well-known subfields of the LIS field. We therefore decided to use this solution. The three clusters can roughly be interpreted as follows. The largest cluster (27 journals) deals with library science, the smallest cluster (7 journals) deals with scientometrics, and the third cluster (14 journals) deals with general information science topics.

We first consider the case of a single integrated LIS field. The recursive MNCS indicator is said to have converged for a certain α if there is virtually no difference between values of the αth-order MNCS indicator and values of the (α + 1)th-order MNCS indicator. For our data, convergence of the recursive MNCS indicator can be observed for α = 20. In our analysis, our main focus therefore is on comparing the first-order MNCS indicator (i.e., the ordinary non-recursive MNCS indicator) with the 20th-order MNCS indicator.

In the case of the first-order MNCS indicator, the top 10 consists of journals from all three subfields. However, journals from the information science and scientometrics subfields seem to slightly dominate journals from the library science subfield.

Let’s now turn to the top 10 journals according to the 20th-order MNCS indicator. This top 10 provides a much more extreme picture. The top 10 is now almost completely dominated by information science and scientometrics journals. There is only one library science journal left, at rank 9.

The top 10 institutes according to both the first-order MNCS indicator and the 20th-order MNCS indicator are listed in Table 5. Comparing the results of the two MNCS indicators, it is clear that institutes which are mainly active in the scientometrics subfield benefit a lot from the use of a higher-order MNCS indicator.

In Table 6, the top 10 journals according to both the first-order MNCS indicator and the 20th-order MNCS indicator is shown. ...  Comparing Table 6 with Table 4, it can be seen that library science journals now play a much more prominent role, both in the case of the first-order MNCS indicator and in the case of the 20th-order MNCS indicator. As a consequence, the top 10 journals now looks much more balanced for both MNCS indicators.

2014年1月24日 星期五

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

本研究探討在地理上的生產力遷移(shifts in productivity)是否發生在書目計量學(bibliometrics)、資訊計量學(informetrics)和科學計量學(scientometrics)等計量學(metrics)領域,也就是歐洲的貢獻明顯地成長,並且北美的貢獻相對來說有減少的情形。

有關計量學的研究,Hood and Wilson (2001)和Stock and Weber(2006)等研究都分析了這個領域的文獻成長情形。Hood and Wilson (2001)回顧了計量學領域的發展,並且比較bibliometrics、scientometrics和informetrics的相關文獻,發現bibliometrics還是在相關領域上使用最廣泛的詞語。Stock and Weber(2006)從觀察中確認這個領域從1980年後便持續地成長。Wolfram (2008)則發現在計量學領域中,北美的文獻有明顯地減少而歐洲則是急遽地增加的情形。

本研究利用bibliometrics、scientometrics、informetrics、cybermetrics、webometrics、citation analysis、link analysis和citation indexes做為檢索的問句,同時再加上Scientometrics和Journal of Informetrics兩種期刊的論文,從Web of Science資料庫中進行檢索。結果共檢索出4404筆論文資料。

在這些論文資料裡,共有75個國家。以地區來區分,歐洲在每個時段上具有最大的貢獻,不論是數量或所占比率都有成長,亞洲所佔的相對比例在22年間有很大的成長,北美雖然在數量上有成長,可是相對的比例呈現緩慢的下降。每個地區的作者會偏好在本身地區的期刊上發表,舉例而言,歐洲作者發表論文的前五個期刊中有四個歐洲期刊,南美也有類似的情形,但是亞洲的情形例外,前五個期刊中有四個是歐洲期刊,另一個則是北美的期刊。

自1990年代中期後,國家間的合作情形增加許多,之前國際合作的論文每年為1到19篇,2008年已大幅增加為96篇。美國是國際合作佔最多的國家,但以地區來說,歐洲平均每個國家的國際合作數為5.78篇論文,多於世界其他部分的4.47篇論文。

此外,歐洲則有許多具有國際合作經驗的機構,共有16所研究機構有國際合作經驗,北美則有8所,亞洲有1所。機構間的合作來說,在1987年每篇論文平均只有1.1個機構,但在2007年則增加為1.96。

本研究且利用MDS、VOSviewer和Pajek將這些論文上的國家與機構之間的合作關係,呈現為圖形。

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).


This investigation was prompted by interest in whether shifts in productivity based on geography are observed in the bibliometrics, informetrics and scientometrics areas.

One of the authors conducted a pilot study to determine whether there have been clear declines in North American contributions to the metrics literature base (Wolfram, 2008). The author found that there was indeed a notable relative decline in North American contributions and a sharp increase in European contributions.

Hood and Wilson (2001) examined the growth of literature of the metrics area. They provided an historical treatment of the development of these areas that included earlier studies of the field. In their research, literature associated with bibliometrics, informetrics and scientometrics was compared for the period 1968–2000. The authors noted that bibliometrics was still the most widely used term for metrics research.

More recently, Stock and Weber(2006) conducted a Web of Science search for records specifically including metrics terms and allied areas. They observed contributions had grown substantially since 1980.

Search parameters included the Boolean ORed result of bibliometrics, scientometrics, informetrics, cybermetrics and webometrics, in truncated form (e.g., webometri*), along with the phrases “citation analysis”, “link analysis” and “citation indexes”. ... These search results were ORed with the two primary journals that publish metrics research that are indexed by WoS, namely Scientometrics and the Journal of Informetrics.

A pair-wise comparison of all collaborations at the national and institutional levels was then conducted from which a cooccurrence matrix could be compiled.

Multidimensional scaling (MDS) analysis was used to visualize the relationships among countries. Because the data represent a type of similarity measure represented as a symmetric matrix, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006).

The recently developed visualization tool VOSviewer (van Eck &Waltman, 2010) was also used to provide an alternate visualization of the relationship outcomes. Like MDS, VOSviewer (http://www.vosviewer.com/) relies on a distance-based approach to mapping informetric relationships. Instead of using more traditional similarity measures to produce a normalized outcome for co-occurrences as used in MDS, relationships are based on association strengths, so the algorithm is somewhat different than PROXSCAL and, therefore, can produce different outcomes. Details of the comparison of different measures can be found in van Eck and Waltman (2009).

The network visualization software Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/) was used as well. Unlike the distance-based mapping of PROXSCAL and VOSviewer, Pajek produces directed or undirected network maps, with the strength of the relationships represented by the thickness of connecting lines between vertices on the map. Distances are used more for clarification, but proximities do not necessarily indicate a stronger relationship.

The search parameters retrieved 4404 publications.

Europe shows the highest levels of contribution, both in absolute and relative terms over the time period of the study. Growth patterns in absolute terms are nonlinear based on trend line analysis in MS Excel; however, the R-squared goodness-of-fit values for even the best fitting models (higher order polynomials) were never more than 0.95, indicating a less than desirable fit.

Relative contributions based on geographic divisions have been largely stable. An exception is Asia, which had an increasing relative contribution over the 22-year time frame of the study. Although North American contributions have continued to increase in absolute numbers, the relative contribution shows a slow average decline over time.

The top five journals listed for each continent demonstrated a regional preference for publication outlets from that region. So, for example, four of the top five journals for European publications were published in Europe, and four of the top five journal outlets for South America were South American. The exception to this was Asia. Four of the top five journals for Asian publications were European and one was North American. This outcome may be a reflection of the data extraction method, the indexing practices of WoS, or a preference during the study time frame for Asian scholars to publish in Western journals.

Seventy-five countries were represented in the record set.

The number of metrics papers published annually that represent collaborations between two or more countries has increased greatly since the mid-1990s. Prior to this time, the number of internationally collaborative papers ranged from 1 to 19 papers annually. Over the last decade this number has increased to a high of 96 papers in 2008.

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).

Sixteen of the institutions on the list are European, eight are North American, and one is Asian. The United States has the largest number of institutions represented (five), followed by Belgium (four – note: one institution merged with another institution to form a new entity).

There has been steady growth in inter-institutional collaboration over the 22 years. The mean number of collaborative institutional partners within the dataset has steadily increased from a low mean of 1.1 institutions per publication in 1987 to a high of 1.96 institutions per publication in 2007.

Europe, and in particular Western Europe, clearly dominates in the production of metrics literature. The United States continues to be the largest singular contributor, but this appears to be changing. North American contributions as a whole continue to increase, but represent a smaller percentage of worldwide production. European contributions have grown tremendously, especially during the last 5 years of the study period. This same period is marked by impressive growth from Asia.

It should be noted that WoS increased its coverage in 2008 by including more regional journals. These inclusions possibly could contribute to the increase in Asian contributions, but the observed growth for Asia was already evident prior to any such additions.

International and inter-institutional collaborations do not necessarily reveal strong geographic affinities, although the multiple institutional affiliations by a number of scholars associated with Flemish institutions do contribute to the strengthening of regional ties. Undoubtedly, the growth of the Internet and increasing availability of other telecommunication technologies have made these collaborations less distance dependent.

2013年12月19日 星期四

van Eck, N. J., Waltman, L., Dekker, R. & van den Berg, J. (2010). A comparison of two techniques for bibliometric mapping: Multidimensional scaling and VOS. Journal of American Society for Information Science and Technology, 61(12), 2045-2061.

van Eck, N. J., Waltman, L., Dekker, R. & van den Berg, J. (2010). A comparison of two techniques for bibliometric mapping: Multidimensional scaling and VOS. Journal of American Society for Information Science and Technology, 61(12), 2045-2061.

本論文從學理與實驗兩方面比較MDS(Multidimesional Scaling)和VOS兩種資訊視覺化的維度縮減方法。作者認為從理論上來看,VOS是在計算Stress Function時以項目(items)間的相似程度(similarity)做為它們在圖形上對應點的接近程度(proximity)加權的一種特殊MDS方法,當兩個項目之間愈相似,它們映射在圖形上的點的接近程度便應該有愈大的加權。由於在實際的資訊視覺化應用上,大多數的項目之間沒有關聯,它們之間的相似程度經常被視為0,傳統的MDS方法也會運用這些項目之間的相似程度進行,導致映射的資料點形成一個接近於圓的圖形,觀測次數較多的項目較有可能會被映射到圓的中心,作者認為這樣的結果是失真的。作者以資訊科學(information science)的作者共被引關係等四種書目計量資料分別進行實驗,結果MDS的兩種相似程度的視覺化結果都相當接近圓,而VOS的結果比較令人滿意。
MDS has been widely used for constructing maps of authors (e.g., McCain, 1990; White & Griffith, 1981; White & McCain, 1998), documents (e.g., Griffith, Small, Stonehill, & Dey, 1974; Small & Garfield, 1985; Small, Sweeney, & Greenlee, 1985), journals (e.g., McCain, 1991), and keywords (e.g., Peters & Van Raan, 1993a, 1993b; Tijssen & Van Raan, 1989).
To determine similarities between items, co-occurrence frequencies are usually transformed using a similarity measure. Two types of similarity measures can be distinguished.
Direct similarity measures (Van Eck & Waltman, 2009; also known as local similarity measures, see Ahlgren, Jarneving, & Rousseau, 2003) determine the similarity between two items by applying a normalization to the co-occurrence frequency of the items. ... Various direct similarity measures are being used in the literature. Especially the cosine and the Jaccard index are very popular. ... We argued that the most appropriate measure for normalizing co-occurrence frequencies is the so-called association strength (e.g., Van Eck & Waltman, 2007b; Van Eck et al., 2006). This measure is also known as the proximity index (e.g., Peters & Van Raan, 1993a; Rip & Courtial, 1984) or as the probabilistic affinity index (e.g., Zitt, Bassecoulard, & Okubo, 2000).
Indirect similarity measures (also known as global similarity measures), on the other hand, determine the similarity between two items by comparing two vectors of co-occurrence frequencies. ... For a long time, the Pearson correlation has been the most popular indirect similarity measure in the literature (e.g., McCain, 1990, 1991; White & Griffith, 1981; White & McCain, 1998). Nowadays, however, it is well known that the Pearson correlation has some undesirable properties (Ahlgren et al., 2003; Van Eck & Waltman, 2008). A well-known indirect similarity measure that does not have these undesirable properties is the cosine.
The aim of MDS is to locate items in a low-dimensional space in such a way that the distance between any two items reflects the similarity or relatedness of the items as accurately as possible. The stronger the relation between two items, the smaller the distance between the items.
To determine the locations of items in a map, MDS minimizes a so-called stress function.  ... MDS determines the locations of items in a map by minimizing the (weighted) sum of the squared differences between on the one hand the transformed proximities of items and on the other hand the distances between items in the map.
Depending on the transformation function f, different types of MDS can be distinguished. The three most important types of MDS are ratio MDS, interval MDS, and ordinal MDS. Ratio and interval MDS are also referred to as metric MDS, while ordinal MDS is also referred to as non-metric MDS. Ratio MDS treats the proximities pij as measurements on a ratio scale. Likewise, interval and ordinal MDS treat the proximities pij as measurements on, respectively, an interval and an ordinal scale. In ratio MDS, f is a linear function without an intercept. In interval MDS, fcan be any linear function, and in ordinal MDS, f can be any monotone function.
The stress function in Equation 3 can be minimized using an iterative algorithm. Various different algorithms are available. A popular algorithm is the SMACOF algorithm (e.g., Borg & Groenen, 2005). This algorithm relies on a technique known
as iterative majorization.
The idea of VOS is to minimize a weighted sum of the squared distances between all pairs of items. The squared distance between a pair of items is weighed by the similarity between the items. To avoid trivial solutions in which all items have the same location, the constraint is imposed that the average distance between two items must be equal to one.
Under certain conditions, MDS and VOS are closely related. More specifically, the proposition indicates that VOS can be regarded as a kind of weighted MDS with proximities and weights chosen in a special way.
For each of the three data sets that we consider, three maps were constructed, one using the MDS-AS (direct similarity measure) approach, one using the MDS-COS (indirect similarity measure) approach, and one using the VOS (direct similarity measure) approach.
Experiment I: 405 authors publishing papers in 36 journals closely related to the Journal of the American Society for Information Science and Technology between 1999 and 2008.
Experiment II: 2079 journals that belong to at least one social science subject category.
Experiment III: 831 keywords that were automatically identified in the abstracts (and titles) of 7492 articles published in 15 operations research journals between 2001 and 2006.
A notable property of the maps produced by the two MDS approaches is that important items (i.e., items with a large number of co-occurrences) tend to be located toward the center of a map. This is especially clear in the case of the authors and keywords data sets. Many relatively unimportant items are scattered throughout the periphery of a map.
The VOS approach seems to produce maps in which important and less important items are distributed fairly evenly over the central and peripheral areas.
In various studies of the field of information science (e.g., Åström, 2007; White & McCain, 1998; Zhao & Strotmann, 2008a,b,c), it has been found that the field consists of two quite independent subfields. We adopt the terminology of Åström (2007) and refer to the subfields as information seeking and retrieval (ISR) and informetrics.
A distinction is sometimes made between ―hard‖ and ―soft‖ ISR research (e.g., Åström, 2007; Persson, 1994; White & McCain, 1998). Hard ISR research is system-oriented and is for example concerned with the development and the experimental evaluation of information retrieval algorithms. Soft ISR research, on the other hand, is user-oriented and studies for example users’ information needs and information behavior. The distinction between hard and soft ISR research is visible in all three maps.
Both (MDS) approaches have a tendency to locate the most prominent authors in the center of a map and less prominent authors in the periphery. Due to this tendency, the separation of subfields becomes more difficult to see.
In these maps(MDS-AS and MDS-COS), a number of prominent ISR authors (e.g., Spink, Wang, and Wilson) are located equally close or even closer to various informetrics authors than to some of their less prominent ISR colleagues. However, contrary to what the maps seem to suggest, there is in fact very little interaction between the prominent ISR authors and the informetrics authors. The relatively small distance between these two groups of authors therefore does not properly reflect the structure of the field of information science. The small distance is merely a technical artifact, caused by the tendency of the MDS-AS and MDS-COS approaches to locate important items in the center of a map.It follows from this observation that distances in maps constructed using the MDS approaches may not always give an accurate representation of the relatedness of items. Hence, in the case of the MDS approaches, the validity of the interpretation of a distance as an (inverse) measure of relatedness seems questionable.
The VOS map in ... does properly reflect the large separation between the prominent ISR authors and the informetrics authors. In this map, the interpretation of a distance as a measure of relatedness therefore seems valid.
This means that in the MDS-AS approach MDS is typically applied to similarity data that consists largely of zeros. MDS attempts to determine the locations of items in a map in such a way that for each pair of items with a similarity of zero the distance between the items is the same. In the case of similarity data that consists largely of zeros, it is not possible to construct a low-dimensional map with exactly the same distance between each pair of items with a similarity of zero. MDS can only try to approximate such a map as closely as possible. Our experiments indicate that the best possible approximation is a map with an almost perfectly circular structure.
From this point of view, one can say that the VOS approach distinguishes itself from the MDS-AS approach in that it does not give equal weight to all pairs of items. The VOS approach gives more weight to more similar pairs of items. It gives little weight to pairs of items with a low similarity. As mentioned above, similarity data is typically dominated by low values, in particular by zeros. ... In the case of the VOS approach, however, pairs of items with a low similarity receive little weight and therefore have little effect on a map. Because of this, the VOS approach does not produce circular maps.

2013年4月29日 星期一

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

information visualization

詞語地圖(term map)能夠將科學領域的結構視覺化。在這裡,詞語指的是能夠代表領域特定概念的詞(words)或片語(phrase),詞語地圖便是為了呈現出領域內重要的詞語之間的關係所產生的圖形。為了製作詞語地圖,本研究提出自動詞語確認(automatic term identification)的方法,以減少專家勞力並避免主觀判斷帶來的問題。考慮到從語料庫確認的詞語必須同時具有單元完整性(unithood)和主題相關性(termhood)兩方面的特質 (Kageura and Umino, 1996),本研究建議的方法包括三個階段:第一階段利用詞類標示器(part-of-speech tagger)產生出來的結果(Schmid,  1994; Schmid, 1995),抽取輸入語料內的名詞片語,做為候選詞語。第二階段比較候選詞語的出現頻率和候選詞語內的第一個詞與其餘部分的出現頻率,計算概似比(likelihood ratio) (Dunning, 1993),評估它們為完整語意單位(semantic unit)的程度,挑選單元完整性比較高的候選詞語。第三階段計算詞語的主題相關性是本研究的重要貢獻。在確認單元完整並與主題相關的詞語後,以每一對詞語之間相關強度(association strength) (Van Eck and Waltman 2009)的值代表它們之間的關係,利用VOS技術 (Van Eck and Waltman 2007a)產生詞語地圖。
本研究建議利用詞語在各主題上的分布傾向來估計每一個詞語的主題相關性。在本研究裡具有較高主題相關性的詞語是只與某一個或較少數主題有較強的關連的詞語。所以對於每一個詞語,本研究建議比較此一詞語在各主題上的分布情形與原先各主題的分布情形,如果差異較大便表示該詞語的主題相關性較高,也就是具有較高主題相關性的詞語對於少數的主題具有區辨力(discriminatory)。但由於每一篇文件都可能包括多個主題,無法單純地統計詞語在各主題上的分布情形以及各主題的分布情形,因此本研究利用機率式隱含語意分析(probabilistic latent semantic analysis, PLSA)的方式(Hofmann, 2001)估計各種詞語在各主題上的分布情形,並與原先各主題的分布情形相比較,找出分布偏向於少數主題的詞語。
評估自動化詞語確認的結果相當困難(Pazienza et al., 2005)。本研究為了評估詞語確認的結果,以15種 ISI主題分類為作業研究(operational research)的期刊,建立了該領域的詞語地圖並且以兩種方式進行評估:第一種方式是比較這種方法與沒有使用PLSA的詞語確認和利用詞語的出現頻率(frequency of occurrence)選取詞語等其他兩種方法的回收率(recall)與精確率(precision)。第二種方式則是由作業研究領域的專家對產生出來的詞語地圖進行品質審核。第一種評估方法的結果顯示除了在最高和最低的回收率以外,本研究建議的方法都比其他兩種方法能夠得到更高的精確率。在專家審查的結果則發現本研究產生的詞語地圖能夠表現出作業研究領域可分為以方法論為導向(methodology-oriented)及以應用為導向(application-oriented)的兩類研究主題,這個結果相當符合專家的想法。但是目前的結果也呈現出這個方法獲得意義較為廣泛的詞語、圖形上沒有包括某些主題以及有些主題非常相近的詞語在圖形上彼此間並不靠近等問題。

A term map is a map that visualizes the structure of a scientific field by showing the relations between important terms in the field.

To evaluate the proposed methodology, we use it to construct a term map of the field of operations research. The quality of the map is assessed by a number of operations research experts.

Other maps show relations between words or keywords based on co-occurrence data (e.g., Rip and Courtial 1984; Peters and Van Raan 1993; Kopcsa and Schiebel 1998; Noyons 1999; Ding et al. 2001). The latter maps are usually referred to as co-word maps.

By a term we mean a word or a phrase that refers to a domain-specific concept. Term maps are similar to co-word maps except that they may contain any type of term instead of only single-word terms or only keywords.

Selection of terms based on their frequency of occurrence in a corpus of documents typically yields many words and phrases with little or no domain-specific meaning. Inclusion of such words and phrases in a term map is highly undesirable for two reasons. First, these words and phrases divert attention from what is really important in the map. Second and even more problematic, these words and phrases may distort the entire structure shown in the map.

However, manual term selection has serious disadvantages as well. The most important disadvantage is that it involves a lot of subjectivity, which may introduce significant biases in a term map. Another disadvantage is that it can be very labor-intensive.

Given a corpus of documents, we first identify the main topics in the corpus. This is done using a technique called probabilistic latent semantic analysis (Hofmann 2001). Given the main topics, we then identify in the corpus the words and phrases that are strongly associated with only one or only a few topics. These words and phrases are selected as the terms to be included in a term map.

An important property of the proposed methodology is that it identifies terms that are not only domain-specific but that also have a high discriminatory power within the domain of interest. This is important because terms with a high discriminatory power are essential for visualizing the structure of a scientific field.

We define unithood as the degree to which a phrase constitutes a semantic unit. Our idea of a semantic unit is similar to that of a collocation (Manning and Schu¨tze 1999). Hence, a semantic unit is a phrase consisting of words that are conventionally used together. The meaning of the phrase typically cannot be fully predicted from the meaning of the individual words within the phrase.

We define termhood as the degree to which a semantic unit represents a domain-specific concept.

Linguistic approaches are mainly used to identify phrases that, based on their syntactic form, can serve as candidate terms.

Statistical approaches are used to measure the unithood and termhood of phrases.

Most terms have the syntactic form of a noun phrase (Justeson and Katz 1995; Kageura and Umino 1996). Linguistic approaches to automatic term identification typically rely on this property. These approaches identify candidate terms using a linguistic filter that checks whether a sequence of words conforms to some syntactic pattern. Different researchers use different syntactic patterns for their linguistic filters (e.g., Bourigault 1992; Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995; Frantzi et al. 2000).

Statistical approaches to measure unithood are discussed extensively by Manning and Schu¨tze (1999). The simplest approach uses frequency of occurrence as a measure of unithood (e.g., Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995). More advanced approaches use measures based on, for example, (pointwise) mutual information (e.g., Church and Hanks 1990; Damerau 1993; Daille et al. 1994) or a likelihood ratio (e.g., Dunning 1993; Daille et al. 1994). Another statistical approach to measure unithood is the C-value (Frantzi et al. 2000). The NC-value (Frantzi et al. 2000) and the SNC-value (Maynard and Ananiadou 2000) are extensions of the C-value that measure not only unithood but also termhood. Other statistical approaches to measure termhood can be found in the work of, for example, Drouin (2003) and Matsuo and Ishizuka (2004). In the field of machine learning, an interesting statistical approach to measure both unithood and termhood is proposed by Wang et al. (2007).

Termhood is measured as the degree to which the occurrences of a semantic unit are biased towards one or more topics.

In the first step of our methodology, we use a linguistic filter to identify noun phrases. We first assign to each word occurrence in the corpus a part-of-speech tag, such as noun, verb, or adjective. The appropriate part-of-speech tag for a word occurrence is determined using a part-of-speech tagger developed by Schmid (1994, 1995). We use this tagger because it has a good performance and because it is freely available for research purposes.

The most common approach to measure unithood is to determine whether a phrase occurs more frequently than would be expected based on the frequency of occurrence of the individual words within the phrase.

To measure the unithood of a noun phrase, we first count the number of occurrences of the phrase, the number of occurrences of the phrase without the first word, and the number of occurrences of the first word of the phrase. In a similar way as Dunning (1993), we then use a so-called likelihood ratio to compare the first number with the last two numbers.

The main idea of the third step of our methodology is to measure the termhood of a semantic unit as the degree to which the occurrences of the unit are biased towards one or more topics.

To measure the degree to which the occurrences of semantic unit uk, where k (belongs to) {1,…,K}, are biased towards one or more topics, we use two probability distributions, namely the distribution of semantic unit uk over the set of all topics and the distribution of all semantic units together over the set of all topics. These distributions are denoted by, respectively, P(tj | uk) and P(tj), where j (belongs to) {1,…, J}. ... The dissimilarity between the two distributions indicates the degree to which the occurrences of uk are biased towards one or more topics. We use the dissimilarity between the two distributions to measure the termhood of uk.

For example, if the two distributions are identical, the occurrences of uk are unbiased and uk most probably does not represent a domain-specific concept. If, on the other hand, the two distributions are very dissimilar, the occurrences of uk are strongly biased and uk is very likely to represent a domain-specific concept.

The dissimilarity between two probability distributions can be measured in many different ways. One may use, for example, the Kullback–Leibler divergence, the Jensen–Shannon divergence, or a chi-square value.

In (3), termhood (uk) is calculated as the negative entropy of this distribution. Notice that termhood (uk) is maximal if P(tj | uk) = 1 for some j and that it is minimal if P(tj | uk) = P(tj) for all j. In other words, termhood (uk) is maximal if the occurrences of uk are completely biased towards a single topic, and termhood (uk) is minimal if the occurrences of uk do not have a bias towards any topic.

In order to allow for a many-to-many relationship between corpus segments and topics, we make use of probabilistic latent semantic analysis (PLSA) (Hofmann 2001).

It was originally introduced as a probabilistic model that relates occurrences of words in documents to so-called latent classes. In the present context, we are dealing with semantic units and corpus segments instead of words and documents, and we interpret the latent classes as topics.

PLSA assumes that each occurrence of a semantic unit in a corpus segment is independently generated according to the following probabilistic process. First, a topic t is drawn from a probability distribution P(tj), where j (belongs to) {1,…,J}. Next, given t, a corpus segment s and a semantic unit u are independently drawn from, respectively, the conditional probability distributions P(si | t), where i (belongs to) {1,…,I}, and P(uk | t), where k (belongs to) {1,…,K}. This then results in the occurrence of u in s.

P(si, uk) =  sum (from j=1 to J) P(tj)P(si|tj)P(uk|tj)

We estimate these parameters using data from the corpus. Estimation is based on the criterion of maximum likelihood. The log-likelihood function to be maximized is given by
L = sum(from i=1 to I) sum(from k=1 to K) nik log P(si, uk)
We use the EM algorithm discussed by Hofmann (1999, Sect. 3.2) to perform the maximization of this function.

After estimating the parameters of PLSA, we apply Bayes’ theorem to obtain a probability distribution over the topics conditional on a semantic unit. This distribution is given by

P(tj|uk) = P(tj)P(uk|tj) / sum (from j=1 to J) (P(P(tj)P(uk|tj))

In a similar way as discussed earlier, we use the dissimilarity between the distributions P(tj | uk) and P(tj) to measure the termhood of uk.

We first selected a number of OR journals. This was done based on the subject categories of Thomson Reuters. The OR field is covered by the category Operations Research & Management Science. Since we wanted to focus on the core of the field, we selected only a subset of the journals in this category. More specifically, a journal was selected if it belongs to the category Operations Research & Management Science and possibly also to the closely related category Management and if it does not belong to any other category. This yielded 15 journals, which are listed in the first column of Table 1.

In the first step of our methodology, the linguistic filter identified 2662 different noun phrases. In the second step, the unithood of these noun phrases was measured. 203 noun phrases turned out to have a rather low unithood and therefore could not be regarded as semantic units. ... The other 2459 noun phrases had a sufficiently high unithood to be regarded as semantic units.

In the third and final step of our methodology, the termhood of these semantic units was measured. To do so, each title-abstract pair in the corpus was treated as a separate corpus segment. For each combination of a semantic unit uk and a corpus segment si, it was determined whether uk occurs in si (nik = 1) or not (nik = 0). Topics were identified using PLSA. This required the choice of the number of topics J. Results for various numbers of topics were examined and compared. Based on our own knowledge of the OR field, we decided to work with J = 10 topics.

The evaluation of a methodology for automatic term identification is a difficult issue. There is no generally accepted standard for how evaluation should be done. We refer to Pazienza et al. (2005) for a discussion of the various problems.

We first perform an evaluation based on the well-known notions of precision and recall. We then perform a second evaluation by constructing a term map and asking experts to assess the quality of this map.

Precision is the number of correctly identified terms divided by the total number of identified terms.

Recall is the number of correctly identified terms divided by the total number of correct terms.

Unfortunately, because the total number of correct terms in the OR field is unknown, we could not calculate the true recall. This is a well-known problem in the context of automatic term identification (Pazienza et al. 2005).

To circumvent this problem, we defined recall in a slightly different way, namely as the number of correctly identified terms divided by the total number of correct terms within the set of all semantic units identified in the second step of our methodology. Recall calculated according to this definition provides an upper bound on the true recall. However, even using this definition of recall, the calculation of precision and recall remained problematic. The problem was that it is very time-consuming to manually determine which of the 2459 semantic units identified in the second step of our methodology are correct terms and which are not. We solved this problem by estimating precision and recall based on a random sample of 250 semantic units.

It is clear from the figure that our methodology outperforms the two simple alternatives. Except for very low and very high levels of recall, our methodology always has a considerably higher precision than the variant of our methodology that does not make use of PLSA.

A term map is a map, usually in two dimensions, that shows the relations between important terms in a scientific field. Terms are located in a term map in such a way that the proximity of two terms reflects their relatedness as closely as possible. That is, the smaller the distance between two terms, the stronger their relation. The aim of a term map usually is to visualize the structure of a scientific field.

It turned out that, out of the 2459 semantic units identified in the second step of our methodology, 831 had the highest possible termhood value. This means that, according to our methodology, 831 semantic units are associated exclusively with a single topic within the OR field. We decided to select these 831 semantic units as the terms to be included in the term map. This yielded a coverage of 97.0%, which means that 97.0% of the title-abstract pairs in the corpus contain at least one of the 831 terms to be included in the term map.

The term map of the OR field was constructed using a procedure similar to the one used in our earlier work (Van Eck and Waltman 2007b). This procedure relies on the association strength measure (Van Eck and Waltman 2009) to determine the relatedness of two terms, and it uses the VOS technique (Van Eck and Waltman 2007a) to determine the locations of terms in the map.

The most serious criticism on the results of the automatic term identification concerned the presence of a number of rather general terms in the map.

Another point of criticism concerned the underrepresentation of certain topics in the term map. There were three experts who raised this issue. One expert felt that the topic of supply chain management is underrepresented in the map. Another expert stated that he had expected the topic of transportation to be more visible. The third expert believed that the topics of combinatorial optimization, revenue management, and transportation are underrepresented.

As discussed earlier, when we were putting together the corpus, we wanted to focus on the core of the OR field and we therefore only included documents from a relatively small number of journals. This may for example explain why the topic of transportation is not clearly visible in the map.

When asked to divide the OR field into a number of smaller subfields, most experts indicated that there are two natural ways to make such a division. On the one hand, a division can be made based on the methodology that is being used, such as decision theory, game theory, mathematical programming, or stochastic modeling. On the other hand, a division can be made based on the area of application, such as inventory control, production planning, supply chain management, or transportation. There were two experts who noted that the term map seems to mix up both divisions of the OR field. According to these experts, one part of the map is based on the methodology-oriented division of the field, while the other part is based on the application-oriented division.

The experts pointed out that sometimes closely related terms are not located very close to each other in the map. One of the experts gave the terms inventory and inventory cost as an example of this problem. In many cases, a problem such as this is probably caused by the limited size of the corpus that was used to construct the map. In other cases, the problem may be due to the inherent limitations of a two-dimensional representation.

Our main contribution consists of a methodology for automatic identification of terms in a corpus of documents. Using this methodology, the process of selecting the terms to be included in a term map can be automated for a large part, thereby making the process less labor-intensive and less dependent on expert judgment. Because less expert judgment is required, the process of term selection also involves less subjectivity.

In general, we are quite satisfied with the results that we have obtained. The precision/recall results clearly indicate that our methodology outperformed two simple alternatives. In addition, the quality of the term map of the OR field constructed using our methodology was assessed quite positively by five experts in the field. However, the term map also revealed a shortcoming of our methodology, namely the incorrect identification of a number of general noun phrases as terms.

As scientific fields tend to overlap more and more and disciplinary boundaries become more and more blurred, finding an expert who has a good overview of an entire domain becomes more and more difficult. This poses serious difficulties for any bibliometric method that relies on expert knowledge.