顯示具有 term identification 標籤的文章。 顯示所有文章
顯示具有 term identification 標籤的文章。 顯示所有文章

2014年2月27日 星期四

Chen, C. M., & Paul, R. J. (2001). Visualizing a knowledge domain's intellectual structure. Computer, 34(3), 65-71.

Chen, C. M., & Paul, R. J. (2001). Visualizing a knowledge domain's intellectual structure. Computer, 34(3), 65-71.
vis_paper
本論文進行ACA(author citation analysis)的研究,以IEEE Computer Graphics and Applications上發表論文的作者為分析對象,選擇353位被引用5次以上的作者,利用他們之間的共被引資訊建立網路圖,結果共有28,638條連結線。經過尋徑網路尺度(pathfinder network scaling)的處理,保留下355條比較重要的連結線。為了發現電腦圖學與應用的專長(specialties),本論文借鏡於White and McCain(1998)的研究,利用PCA(principal component analysis)方法對共被引資料進行因素分析(factor analysis),結果共得到60個專長,5個較大的專長共可以解釋39%的變異數,而這5個專長分別是Rendering and ray tracing、Computer vision、Geometric modeling and computer-aided design、Volume rendering和Modeling nature。同時也在網路圖上呈現被歸類為這5個專長的作者,來觀察他們在網路圖上的分布情形。
ACA, a special type of citation analysis, focuses on intellectual connections between authors as reflected through the scientific literature. The author co-citation relationship links two authors by how often other authors reference their work together. Author co-citation patterns provide the basis for constructing an alternative view to a knowledge structure.
Pathfinder uses a filtering criterion known as the triangle inequality condition to determine whether to remove or retain each link in the original network. Triangle inequality requires that the length of a path connecting two points in the network should not be longer than the length of other alternative paths connecting the two points, but go through extra intermediate points.
We began by studying author co-citation patterns found in IEEE Computer Graphics and Applications magazine for a period of 18 years. ...  Among them, we entered into the author co-citation analysis only the 353 authors who received more than five citations in CG&A. Although this snapshot derives from a limited viewpoint—the literature of  computer graphics certainly stretches beyond the scope of CG&A— intellectual groupings of these 353 authors provide the basis for visualizing the computer graphics knowledge domain. ... The original author co-citation network contains as many as 28,638 links, which constitutes 46 percent of all possible links, excluding self-citations. Because this many links would clutter visualizations, we applied Pathfinder network scaling to reduce their number to 355.
We enhanced the network by coloring it according to the results generated using principal component analysis (PCA). PCA identified 60 specialties in computer graphics. The largest (rendering and ray tracing) and second-largest (computer vision) accounted for 13 percent and 11 percent of the variance, respectively. The five largest specialties accounted for 39 percent of the variance. Remaining specialties are relatively small.
Factor 1: Rendering and ray tracing.
Factor 2: Computer vision.
Factor 3: Geometric modeling and computer-aided design.
Factor 4: Volume rendering.
Factor 5: Modeling nature.
The knowledge landscape visualizes intellectual structures. A virtual landscape like this provides an intuitive gateway for users to access the scientific literature. Researchers new to a field can gain a useful overview by using the knowledge landscape to establish their own mental model of the field and track the development of their own domain.

2014年2月3日 星期一

Dutt, B., Garg, K. C., & Bali, A. (2003). Scientometrics of the international journal Scientometrics. Scientometrics, 56(1), 81-93.

Dutt, B., Garg, K. C., & Bali, A. (2003). Scientometrics of the international journal Scientometrics. Scientometrics, 56(1), 81-93.

過去關於科學計量領域的計量分析結果:Wouters and Leydesdorff (1994)根據Price指標的分類,指出科學計量學並未成為一門硬性的社會科學(hard social science),Schoepflin and Glänzel (2001)則認為這個領域的異質性很高,每一個次領域都有它本身的特性。

本研究針對以下的問題進行探討:確認Scientometrics期刊上1978到2001年發表論文資料的主題,分析這期間論文的分布情形,不同國家在不同主題上的貢獻,具有主要生產力的機構,並藉由合著關係發掘國內與國際間的合作情形。本研究將論文的主題分為科學計量評估(scientometric assessment)、引用與叢集分析(citation and cluster analysis)、科學計量分布(scientometric distribution)、科學的歷史(history of science)、科學合作(scientific collaboration)、科學計量學的理論研究(theoretical studies on scientometrics),不在上述主題的論文則歸類為其他。

結果發現:論文數最多的主題是科學計量評估(scientometric assessment),這個現象反映出科學政策的制定逐漸運用科學計量工具的事實,其次是理論研究(theoretical studies)。在前期(1978-1986年),科學的歷史(history of science)方面的論文較多,其次是引用與叢集分析(citation and cluster analysis);科學計量分布(scientometric distribution)在前期與中期(1987-1994年)都相當重要,但後期(1994-2001)逐漸減少;後期具有最重要地位的主題則是科學合作(scientific collaboration)。

美國是目前生產力最高的國家,共占17.7%的論文,主要的8個歐洲國家則共佔47.6%,但美國在論文所佔的比例逐年減少,加拿大與前蘇聯有同樣的情形,但荷蘭、印度、法國和日本在上升中;從每一個機構平均發表的論文數可以看出這個領域的生產力相當分散,1317篇論文的作者資料共來自1538個機構,但是有1109篇論文是單一機構發表,兩個或以上的機構發表的論文只有208篇;在1538個機構中,發表超過15篇或以上論文的機構共有8個,匈牙利和荷蘭各有2個,其餘的4個機構分別位於印度、比利時、英國和美國;雖然目前的論文以單一作者為主,論文的平均作者數僅為1.73,但多位作者的論文雖然僅占18%,但正逐漸增加。

The study indicates that the US share of papers is constantly on the decline while that of the Netherlands, India, France and Japan is on the rise.

The research output is highly scattered as indicated by the average number of papers per institution.

The scientometric output is dominated by the single authored papers, however, multi-authored papers are gaining momentum.

However, Wouters and Leydesdorff [1] presented a combined bibliometric and social network analysis of papers published in first 25 volumes of Scientometrics, and concluded that scientometrics has not become a hard social science as reflected by the values of Price Index.

In another study, Schoepflin and Glänzel [2] point out that the field of scientometrics is heterogeneous, and each sub-discipline has its own characteristics.

The objectives of the study are:
(i) to identify the scientometric themes on which papers have been published in volumes 1(1978) to 50 (2001), and to find out as to how the emphasis on different themes have changed during different periods;
(ii) to examine the distribution of output of different countries during 1978 -2001, and to analyse the change in the trend, if any;
(iii) to study the relative research emphasis of different countries on different scientometric themes;
(iv) to identify the most productive institutions, and to study the scientometric themes they have dealt with;
(v) to study the pattern of co-authorship and the pattern of domestic as well as international collaboration.

The entire data set was classified into seven groups:
scientometric assessment;
citation and cluster analysis;
scientometric distribution;
history of science;
scientific collaboration;
theoretical studies on scientometrics.
Papers which could not fit into these categories were kept under ‘others’.

An analysis of the data indicates that about one-third of the papers published in Scientometrics deal with scientometric assessment which mainly include cross-national, national and institutional assessment, besides evaluation of journals, bibliometric performance indicators, funding and performance, and S&T indicators. This was followed by theoretical studies (Table 1).

From the values of the Activity Index presented in Table 1, it is observed that the priorities of different themes kept changing during different periods. For instance, during 1978-1986 ‘history of science’ followed by ‘citation and cluster analysis’ were the areas of maximum emphasis.

Studies dealing with ‘scientometrics distribution’ got almost the same priority during 1978-1986 and 1987-1994, but emphasis on this theme has gone down considerably in the last block.

During 1994-2001, studies dealing with ‘scientific collaboration’ got maximum priority followed by ‘scientomeric assessment’.

Major contribution (>=2%) of the total output came from 13 countries listed in Table 2. The distribution of papers presented in Table 2 indicates that USA tops the list of publications which are 17.7 per cent of the total world output.

The values of the Activity Index for different countries (Table2) indicate that during the last two blocks, i.e. 1987-1994 and 1994-2001, the productivity of the USA has declined considerably. Similar is the case with Canada and the former USSR.

Further analysis of data presented in Table 2 indicates that Scientometrics is getting Euro-centred, as 8 countries of Europe listed in Table 2 have contributed 47.6 per cent of the total output. The share may be greater, if the output from other European countries not listed in Table 2 is included.

The total output of 1317 papers published in 50 volumes of Scientometrics came from 1538 institutions. 1109 papers were published involving only a single institute and the rest 208 involved collaboration either with 2 or more institutes.

Number of such institutes which published 15 or more papers is only 8 and their share in the total output is 259 (19.66 %). Of the 8 prolific institutions, two are from Hungary, two from the Netherlands, and one each from India, Belgium, UK and USA.

The results presented in Table 5 indicate that slightly more than half of the papers were single authored and the rest were written by either two or more authors. The share of multi-authored papers (>=3) is much less (18%) only as compared to single or two authored papers.

A study carried out by Cunningham and Dillon [7] for authorship pattern in library and information science indicates average number of authors per paper for information science is 1.17. In scientometrics the average number of authors per paper is 1.73 which indicates a better collaboration than library and information science.

The values of DCI for Spain, France, India and Japan were much higher than the world average indicating a good domestic collaboration. However, except Spain all these countries had very low values of ICI, which indicates that these countries have a poor international collaboration. On the other hand UK, Hungary, Belgium, Canada and Germany had good international collaboration as reflected by the values of ICI.

The focus of scientometric studies is shifting from the history of science and scientometrics distribution to scientific collaboration and scientometric assessment.

Scientometric assessment constitutes about 34% of the total output of the papers which is the highest among all the themes. Emphasis on scientometric assessment studies reflects the growing realisation of its utility as a tool for science policy making.

2014年1月25日 星期六

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

本研究利用詞語共現分析(co-word analysis)技術,區分出六個圖書資訊學的研究主題:兩個書目計量學主題、一個資訊檢索主題、一個一般議題、一個網路計量學主題以及一個專利研究主題。

詞語共現分析根據詞語共同在文件出現的現象描述文件的內容,利用共同出現的相對強度呈現領域的概念網絡(concept networks)。目前已經有植物生物學(de Looze and Lemarie, 1997) 、凝態物理(Bhattacharya and Basu, 1998)、化學工程(Peters and van Raan, 1993)、資訊檢索(Ding, Chowdhury, and Foo, 2001)以及 醫學(Onyancha and Ocholla, 2005)等多個領域曾利用詞語共現分析技術來研究領域內的概念網絡。Van Raan and Tijssen (1993)討論基於詞語共現分析的書目計量在知識論的潛力(epistemological potentitals)。相較於共被引分析,詞語共現分析能應用在沒有引用索引的資料,而且共被引分析會因為在領域的變動與趨勢以及引用者的行為而變得複雜(Noyons & van Raan, 1998)。雖然Leydesdorff (1997)認為詞語的意義隨它們與其他詞語關係的頻率及其出現位置,會有所改變;但Courtial (1998)則是認為詞語共現分析中的詞語,並非做為用來代表某種意義的語言單位,而僅僅是文本間的連結指標。

本研究列舉幾個應用文字資訊為基礎的書目計量方法在圖書資訊學研究主題分析的研究:Courtial(1994)以詞語共現分析對這個領域進行探討,結果發現這個領域包含傳統圖書館學、資訊檢索、科學計量學、資訊計量學、專利分析以及最近興起的網路計量學。Glänzel及其同事整合全文為基礎的結構分析(full-text based structural analysis)和傳統的書目計量方法探討書目計量學及其次領域(Glenisson, Glänzel, and Persson, 2005; Glenisson, Glänzel, Janssens, and De Moor, 2005; Janssens, Glenisson, Glänzel, and De Moor, 2005)。

本研究所使用的分析技術包括:文本抽取(text extraction)、前處理(preprocessing)、多維度尺度(multidimensional scaling)以及Ward’s階層叢集(Ward's hierarchical clustering),並且利用向量空間模式(vector space model) (Salton & McGill, 1986)和隱藏語意分析(latent semantic analysis) (Deerwester et al., 1990)測量文件間相似程度的估計值。以論文彼此間的相似程度,將論文映射成二維圖形的結果如下,此圖形並且標示出每篇論文的期刊:

Scientometrics的論文主要分布在標示為1與2的兩個橢圓附近,橢圓1的主題為書目計量,橢圓2則為專利分析。橢圓5上的論文主要來自Information Processing and Management和Journal of the American Society for Information Science and Technology,其主題為資訊檢索。橢圓12的論文傾向於社會方面的主題,除了Journal of the American Society for Information Science and Technology以外,還包括Journal of Information Science和Journal of Documentation。正中央標示為14的橢圓,其主題與網路相關,所有的期刊均有這個主題的相關論文。

以Ward's叢集分析將所有論文進行歸類,最佳的結果共分為六個叢集。本研究並且根據每個叢集上論文的重要詞語以及中心的論文給予叢集的名稱。在二維圖形上標示各種叢集的結果如下:

六個叢集可以圖形上的斜線分為兩群,斜線以下為Bibliometrics1、Bibliometrics2和Patent Analysis,以上則為Webometrics、Information Retrieval和Social Aspects,但六個叢集中以Patent Analysis和其他叢集較分離。書目計量相關論文分為兩個叢集:Bibliometrics1和Bibliometrics2。Bibliometrics1與科學裡的合作關係(collaboration in science)、引用分析(citation analyses)和國家研究成效(national research performance)等主題相關,Bibliometrics2則主要為方法學和書目計量理論相關的論文。

為了找出各期刊分別著重的主題,除了比較上面的兩個圖形,另外還將叢集和期刊的關係映射成圖形。結果發現Information Processing and Management和Information Retrieval幾乎重疊,這個現象表示Information Processing and Management上的論文和Information Retrieval十分相關。Social Aspects和Webometrics相當靠近Journal of the American Society for Information Science and Technology、Journal of Information Science和Journal of Documentation三種期刊。事實上,除了Scientometrics以外,Social Aspects和其他期刊的距離大約相等。最後,Scientometrics則是落在Bibliometrics1、Bibliometrics2和Patent Analysis構成的三角形中心。

The optimum solution for clustering LIS is found for six clusters. The combination of different mapping techniques, applied to the full text of scientific publications, results in a characteristic tripod pattern. Besides two clusters in bibliometrics, one cluster in information retrieval and one containing general issues, webometrics and patent studies are identified as small but emerging clusters within LIS.

The method was developed by Callon, Courtial, Turner, and Brain (1983), more than two decades ago, for purposes of evaluating research. The methodological foundation of co-word analysis is the idea that the co-occurrence of words describes the contents of documents. By measuring the relative intensity of these co-occurrences, simplified representations of a field’s concept networks can be illustrated (Callon, Courtial, & Laville, 1991).

Van Raan and Tijssen (1993) have discussed the ‘‘epistemological’’ potentials of bibliometric mapping based on co-word analysis.

Leydesdorff (1997) analysed 18 full-text articles and sectional differences therein, and considered that the subsumption of similar words under keywords assumes stability in the meanings, but that words can change both in terms of frequencies of relations with other words, and in terms of positional meaning from one text to another. This fluidity was expected to destabilize representations of developments of the sciences on the basis of co-occurrences and co-absences of words.

However, Courtial (1998) replied that words, in co-word analysis, are not used as linguistic items to mean something, but as indicators of links between texts.

Many researchers have used this methodology to investigate concept networks in different fields, among others, de Looze and Lemarie (1997) in plant biology, Bhattacharya and Basu (1998) in condensed matter physics, Peters and van Raan (1993) in chemical engineering, Ding, Chowdhury, and Foo (2001) in information retrieval (IR) and Onyancha and Ocholla (2005) in medicine.

The reason why the emphasis has shifted from co-citation analysis to co-word techniques is twofold. The first reason is a practical one; co-word analysis allows application to non-citation indexes as well. The second relates to methodology; co-citation analysis complicates the combined analysis of field dynamics and trends in the actors’ activity (Noyons & van Raan, 1998).

Bonnevie (2003) has used primary bibliometric indicators to analyse the Journal of Information Science, while He and Spink (2002) compared the distribution of foreign authors in Journal of Documentation and Journal of the American Society for Information Science and Technology.

Bibliometric trends of the journal Scientometrics, another important journal of the field, have been examined by Schubert and Maczelka (1993), Wouters and Leydesdorff (1994), Schoepflin and Glänzel (2001), Schubert (2002), Dutt, Garg, and Bali (2003).

The main journals of the field were also analysed in terms of journal co-citation and keyword analyses (Marshakova, 2003; Marshakova-Shaikevich, 2005).

The co-citation network of highly cited authors active in the field of IR was studied by Ding, Chowdhury, and Foo (1999).

Finally, Persson (2000, 2001) analysed author co-citation networks on basis of documents published in the journal Scientometrics.

Courtial (1994) has studied the dynamics of the field by analysing the co-occurrence of words in titles and abstracts. Courtial described scientometrics as a hybrid field consisting of invisible colleges, conditioned by demands on the part of scientific research and end-users. Although this situation might have somewhat changed during the last decade, this conclusion illustrates how heterogeneous the much broader field of LIS – comprising subdisciplines such as traditional library science, IR, scientometrics, informetrics, patent analyses and most recently the emerging specialty of webometrics – nowadays is.

In recent papers, Glenisson, Gla¨nzel, and Persson (2005), Glenisson, Gla¨nzel, Janssens, and De Moor (2005), Janssens, Glenisson, Gla¨nzel, and De Moor (2005) have applied full-text based structural analysis in combination with ‘‘traditional’’ bibliometric methods to bibliometrics and its subdisciplines.

The full-text analysis consisted of text extraction, preprocessing, multidimensional scaling, and Ward’s hierarchical clustering (Jain & Dubes, 1988).

In short, the textual information is encoded in the vector space model using the TF-IDF weighting scheme, and similarities are calculated as the cosine of the angle between the vector representations of two items (see Salton & McGill, 1986; Baeza-Yates & Ribeiro-Neto, 1999).

The term-by-document matrix A is again transformed into a latent semantic index Ak (LSI), an approximation of A, but with rank k much lower than the term or document dimension of A. A latent semantic analysis is advisable, especially when dealing with full-text documents in which a lot of noise is observed.

One advantage of LSI is the fact that synonyms or different term combinations describing the same concept are mapped on the same factor, based on the common context in which they generally appear (Berry et al., 1995; Deerwester et al., 1990).

A lot of time was devoted to the detection of phrases. Since the best phrase candidates can be found in noun phrases, the programs LT POS and LT CHUNK4 have first been applied to detect all noun phrases in the complete document collection.

MDS represents all high-dimensional points (documents) in a two- or three-dimensional space in a way that the pairwise distances between points approximate the original high-dimensional distances as precisely as possible (see Mardia, Kent, & Bibby, 1979).

The agglomerative hierarchical cluster algorithm using Ward’s method (see Jain & Dubes, 1988) was chosen to subdivide the documents into clusters. ... One of the disadvantages of agglomerative hierarchical clustering is that wrong choices (merges) that are made by the algorithm in an early stage can never be repaired (Kaufman & Rousseeuw, 1990). What we sometimes observe when using hierarchical clustering is the forming of one very big cluster and a few small very specific clusters.

The journal Scientometrics can be largely separated from the other journals (which is also confirmed by the different term profile in the table of Appendix 1), and exhibits two different foci (best visible in Fig. 4).



The first ‘‘leg’’, indicated by the ellipse with number 1 and by and large containing the first focus of the journal Scientometrics, clearly contains papers in bibliometrics. The 10 best TF-IDF terms for ‘‘leg’’ #1 are: citat, cite, impact factor, self citat, co citat, scienc citat index, citat rate, isi, countri and bibliometr.

The second ‘‘leg of Scientometrics’’, indicated by number 2, is characterised by the best terms patent, industri, biotechnolog, inventor, invent, compani, firm, thin film, brazilian and citat. The JIS paper (#3) embedded in this patent ‘‘leg’’ might be considered an outlier for that journal, but it was put in the right place since it is concerned with ‘‘The many applications of patent analysis’’ (Appendix 2: Breitzman & Mogee, 2002).

An important focus of LIS is indicated by ellipse #5 and can be profiled as ‘‘Information Retrieval’’ (IR) when looking at the highest scoring terms: queri, search engin, web, node, music, imag, xml, vector and weight.

The fourth distinguishable subpart of LIS (#12) is about digit, internet, servic, seek, behaviour, health, knowledg manag, organiz, social and respond; so encompassing the more social aspects.

The remaining large subpart is somewhat the central part (#14). It consists of papers leading to a mean profile containing the terms web, web site, classif, domain, web page, languag, scientist, region, catalog, and web impact factor.

The term network of Cluster 1 allowed the conclusion that the papers belonging to this cluster are concerned with domain studies, studies of collaboration in science, citation analyses, national research performance and similar issues.



The medoid is a paper by Persson et al. on ‘‘Inflationary bibliometric values: The role of scientific collaboration and the need for relative indicators in evaluative studies’’ (Appendix 2: Persson et al., 2004). This is a methodological paper with strong implications for research evaluation, combining research collaboration with citation analysis and construction of national science indicators.

The smaller bibliometrics cluster (Cluster 3: manually labelled as ‘‘Bibliometrics2’’) is of more methodological/theoretical nature.




The medoid is the state-of-the-art report ‘‘Journal impact measures in bibliometric research’’ (Appendix 2: Gla¨nzel & Moed, 2002).

The term networks for the two bibliometrics clusters just described contain a few overlapping terms (bibliometr, chemistri, citat, citat rate, cite, cluster, countri, impact factor, isi, physic, rank and scienc citat index). The MDS plot of Fig. 15 confirms that there is no clear border between Bibliometrics1 and Bibliometrics2, but that there is a gradual transition.

The almost tiny Cluster 2 (19 papers, Fig. 10) represents patent analysis.


A paper on ‘‘Methods for using patents in cross-country comparisons’’ forms the medoid of this cluster (Appendix 2: Archambault, 2002).

Cluster 4, with 282 papers, is the largest one. We have labelled it ‘‘Information Retrieval’’.


The medoid paper is entitled ‘‘Querying and ranking XML documents’’ (Appendix 2: Schlieder & Meuss, 2002).

Cluster 5, with 62 papers, belongs to the small clusters. Both terms and papers close to the medoid characterise this cluster as ‘‘Webometrics’’.


The medoid paper is entitled ‘‘Motivations for academic web site interlinking: evidence for the Web as a novel source of information on informal scholarly communication’’ (Appendix 2: Wilkinson et al., 2003).

Cluster 6 (213 papers) proved to be the most heterogeneous cluster. We have labelled it ‘‘Social’’, however, we could also have called it ‘‘General & miscellaneous issues’’.



‘‘Approaches to user-based studies in information seeking and retrieval: a Sheffield perspective’’ is the title of the medoid paper (Appendix 2: Beaulieu, 2003).


The Patent cluster can be clearly separated from the rest of LIS. The subspace under the line is almost completely occupied by Bilbiometrics1, Bibliometrics2 and Patent.




IR and IPM almost collide in this 2D projection (Fig. 20). This means that Cluster 4 (‘‘IR’’) is very close to the scope of this journal.

The ‘‘Social’’ cluster with general and miscellaneous topics as well as ‘‘Webometrics’’ are close to JIS, JDoc and JASIST, too. Moreover, the ‘‘Social’’ cluster is almost equidistant to all traditional journals in Information Science.

The remaining three clusters, namely Bibliometrics1, Bibliometrics2 and Patent, form a triangle in the centre of which the journal Scientometrics is located. The relatively large distances among these clusters and between each cluster and the journal, strongly indicate that a quite large spectrum of bibliometric, technometric and informetric research using different vocabularies is covered by the journal Scientometrics. This observation is in line with the findings by Schoepflin and Gla¨nzel (2001) that scientometrics consists of several subdisciplines such as informetric theory, empirical studies, indicator engineering, methodological studies, sociological approach and science policy; and that case studies and methodology became dominant by the late 1990s. At the end of the 1990s, also technology related studies based on patent statistics became an emerging subdiscipline of the field.

We have found two clusters in bibliometrics, of which a big one in applied bibliometrics/research evaluation and a smaller one in methodological/theoretical issues; also we have found two large clusters in information retrieval and general and miscellaneous issues and, finally, two small emerging clusters in webometrics and patent and technology studies. Within the IR cluster, we have found a small subcluster on music retrieval, which might be a temporary phenomenon since the journal JASIST has published a special issue on this topic.

According to the expectation, IR, General issues and Webometrics were represented by four of the five journals, namely JIS, IPM, JASIST and JDoc, while the two bibliometrics and the patent clusters were the domain of the journal Scientometrics.

2013年4月29日 星期一

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

Van Eck, N. J., Waltman, L., Noyons, E. C., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. Scientometrics, 82(3), 581-596.

information visualization

詞語地圖(term map)能夠將科學領域的結構視覺化。在這裡,詞語指的是能夠代表領域特定概念的詞(words)或片語(phrase),詞語地圖便是為了呈現出領域內重要的詞語之間的關係所產生的圖形。為了製作詞語地圖,本研究提出自動詞語確認(automatic term identification)的方法,以減少專家勞力並避免主觀判斷帶來的問題。考慮到從語料庫確認的詞語必須同時具有單元完整性(unithood)和主題相關性(termhood)兩方面的特質 (Kageura and Umino, 1996),本研究建議的方法包括三個階段:第一階段利用詞類標示器(part-of-speech tagger)產生出來的結果(Schmid,  1994; Schmid, 1995),抽取輸入語料內的名詞片語,做為候選詞語。第二階段比較候選詞語的出現頻率和候選詞語內的第一個詞與其餘部分的出現頻率,計算概似比(likelihood ratio) (Dunning, 1993),評估它們為完整語意單位(semantic unit)的程度,挑選單元完整性比較高的候選詞語。第三階段計算詞語的主題相關性是本研究的重要貢獻。在確認單元完整並與主題相關的詞語後,以每一對詞語之間相關強度(association strength) (Van Eck and Waltman 2009)的值代表它們之間的關係,利用VOS技術 (Van Eck and Waltman 2007a)產生詞語地圖。
本研究建議利用詞語在各主題上的分布傾向來估計每一個詞語的主題相關性。在本研究裡具有較高主題相關性的詞語是只與某一個或較少數主題有較強的關連的詞語。所以對於每一個詞語,本研究建議比較此一詞語在各主題上的分布情形與原先各主題的分布情形,如果差異較大便表示該詞語的主題相關性較高,也就是具有較高主題相關性的詞語對於少數的主題具有區辨力(discriminatory)。但由於每一篇文件都可能包括多個主題,無法單純地統計詞語在各主題上的分布情形以及各主題的分布情形,因此本研究利用機率式隱含語意分析(probabilistic latent semantic analysis, PLSA)的方式(Hofmann, 2001)估計各種詞語在各主題上的分布情形,並與原先各主題的分布情形相比較,找出分布偏向於少數主題的詞語。
評估自動化詞語確認的結果相當困難(Pazienza et al., 2005)。本研究為了評估詞語確認的結果,以15種 ISI主題分類為作業研究(operational research)的期刊,建立了該領域的詞語地圖並且以兩種方式進行評估:第一種方式是比較這種方法與沒有使用PLSA的詞語確認和利用詞語的出現頻率(frequency of occurrence)選取詞語等其他兩種方法的回收率(recall)與精確率(precision)。第二種方式則是由作業研究領域的專家對產生出來的詞語地圖進行品質審核。第一種評估方法的結果顯示除了在最高和最低的回收率以外,本研究建議的方法都比其他兩種方法能夠得到更高的精確率。在專家審查的結果則發現本研究產生的詞語地圖能夠表現出作業研究領域可分為以方法論為導向(methodology-oriented)及以應用為導向(application-oriented)的兩類研究主題,這個結果相當符合專家的想法。但是目前的結果也呈現出這個方法獲得意義較為廣泛的詞語、圖形上沒有包括某些主題以及有些主題非常相近的詞語在圖形上彼此間並不靠近等問題。

A term map is a map that visualizes the structure of a scientific field by showing the relations between important terms in the field.

To evaluate the proposed methodology, we use it to construct a term map of the field of operations research. The quality of the map is assessed by a number of operations research experts.

Other maps show relations between words or keywords based on co-occurrence data (e.g., Rip and Courtial 1984; Peters and Van Raan 1993; Kopcsa and Schiebel 1998; Noyons 1999; Ding et al. 2001). The latter maps are usually referred to as co-word maps.

By a term we mean a word or a phrase that refers to a domain-specific concept. Term maps are similar to co-word maps except that they may contain any type of term instead of only single-word terms or only keywords.

Selection of terms based on their frequency of occurrence in a corpus of documents typically yields many words and phrases with little or no domain-specific meaning. Inclusion of such words and phrases in a term map is highly undesirable for two reasons. First, these words and phrases divert attention from what is really important in the map. Second and even more problematic, these words and phrases may distort the entire structure shown in the map.

However, manual term selection has serious disadvantages as well. The most important disadvantage is that it involves a lot of subjectivity, which may introduce significant biases in a term map. Another disadvantage is that it can be very labor-intensive.

Given a corpus of documents, we first identify the main topics in the corpus. This is done using a technique called probabilistic latent semantic analysis (Hofmann 2001). Given the main topics, we then identify in the corpus the words and phrases that are strongly associated with only one or only a few topics. These words and phrases are selected as the terms to be included in a term map.

An important property of the proposed methodology is that it identifies terms that are not only domain-specific but that also have a high discriminatory power within the domain of interest. This is important because terms with a high discriminatory power are essential for visualizing the structure of a scientific field.

We define unithood as the degree to which a phrase constitutes a semantic unit. Our idea of a semantic unit is similar to that of a collocation (Manning and Schu¨tze 1999). Hence, a semantic unit is a phrase consisting of words that are conventionally used together. The meaning of the phrase typically cannot be fully predicted from the meaning of the individual words within the phrase.

We define termhood as the degree to which a semantic unit represents a domain-specific concept.

Linguistic approaches are mainly used to identify phrases that, based on their syntactic form, can serve as candidate terms.

Statistical approaches are used to measure the unithood and termhood of phrases.

Most terms have the syntactic form of a noun phrase (Justeson and Katz 1995; Kageura and Umino 1996). Linguistic approaches to automatic term identification typically rely on this property. These approaches identify candidate terms using a linguistic filter that checks whether a sequence of words conforms to some syntactic pattern. Different researchers use different syntactic patterns for their linguistic filters (e.g., Bourigault 1992; Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995; Frantzi et al. 2000).

Statistical approaches to measure unithood are discussed extensively by Manning and Schu¨tze (1999). The simplest approach uses frequency of occurrence as a measure of unithood (e.g., Dagan and Church 1994; Daille et al. 1994; Justeson and Katz 1995). More advanced approaches use measures based on, for example, (pointwise) mutual information (e.g., Church and Hanks 1990; Damerau 1993; Daille et al. 1994) or a likelihood ratio (e.g., Dunning 1993; Daille et al. 1994). Another statistical approach to measure unithood is the C-value (Frantzi et al. 2000). The NC-value (Frantzi et al. 2000) and the SNC-value (Maynard and Ananiadou 2000) are extensions of the C-value that measure not only unithood but also termhood. Other statistical approaches to measure termhood can be found in the work of, for example, Drouin (2003) and Matsuo and Ishizuka (2004). In the field of machine learning, an interesting statistical approach to measure both unithood and termhood is proposed by Wang et al. (2007).

Termhood is measured as the degree to which the occurrences of a semantic unit are biased towards one or more topics.

In the first step of our methodology, we use a linguistic filter to identify noun phrases. We first assign to each word occurrence in the corpus a part-of-speech tag, such as noun, verb, or adjective. The appropriate part-of-speech tag for a word occurrence is determined using a part-of-speech tagger developed by Schmid (1994, 1995). We use this tagger because it has a good performance and because it is freely available for research purposes.

The most common approach to measure unithood is to determine whether a phrase occurs more frequently than would be expected based on the frequency of occurrence of the individual words within the phrase.

To measure the unithood of a noun phrase, we first count the number of occurrences of the phrase, the number of occurrences of the phrase without the first word, and the number of occurrences of the first word of the phrase. In a similar way as Dunning (1993), we then use a so-called likelihood ratio to compare the first number with the last two numbers.

The main idea of the third step of our methodology is to measure the termhood of a semantic unit as the degree to which the occurrences of the unit are biased towards one or more topics.

To measure the degree to which the occurrences of semantic unit uk, where k (belongs to) {1,…,K}, are biased towards one or more topics, we use two probability distributions, namely the distribution of semantic unit uk over the set of all topics and the distribution of all semantic units together over the set of all topics. These distributions are denoted by, respectively, P(tj | uk) and P(tj), where j (belongs to) {1,…, J}. ... The dissimilarity between the two distributions indicates the degree to which the occurrences of uk are biased towards one or more topics. We use the dissimilarity between the two distributions to measure the termhood of uk.

For example, if the two distributions are identical, the occurrences of uk are unbiased and uk most probably does not represent a domain-specific concept. If, on the other hand, the two distributions are very dissimilar, the occurrences of uk are strongly biased and uk is very likely to represent a domain-specific concept.

The dissimilarity between two probability distributions can be measured in many different ways. One may use, for example, the Kullback–Leibler divergence, the Jensen–Shannon divergence, or a chi-square value.

In (3), termhood (uk) is calculated as the negative entropy of this distribution. Notice that termhood (uk) is maximal if P(tj | uk) = 1 for some j and that it is minimal if P(tj | uk) = P(tj) for all j. In other words, termhood (uk) is maximal if the occurrences of uk are completely biased towards a single topic, and termhood (uk) is minimal if the occurrences of uk do not have a bias towards any topic.

In order to allow for a many-to-many relationship between corpus segments and topics, we make use of probabilistic latent semantic analysis (PLSA) (Hofmann 2001).

It was originally introduced as a probabilistic model that relates occurrences of words in documents to so-called latent classes. In the present context, we are dealing with semantic units and corpus segments instead of words and documents, and we interpret the latent classes as topics.

PLSA assumes that each occurrence of a semantic unit in a corpus segment is independently generated according to the following probabilistic process. First, a topic t is drawn from a probability distribution P(tj), where j (belongs to) {1,…,J}. Next, given t, a corpus segment s and a semantic unit u are independently drawn from, respectively, the conditional probability distributions P(si | t), where i (belongs to) {1,…,I}, and P(uk | t), where k (belongs to) {1,…,K}. This then results in the occurrence of u in s.

P(si, uk) =  sum (from j=1 to J) P(tj)P(si|tj)P(uk|tj)

We estimate these parameters using data from the corpus. Estimation is based on the criterion of maximum likelihood. The log-likelihood function to be maximized is given by
L = sum(from i=1 to I) sum(from k=1 to K) nik log P(si, uk)
We use the EM algorithm discussed by Hofmann (1999, Sect. 3.2) to perform the maximization of this function.

After estimating the parameters of PLSA, we apply Bayes’ theorem to obtain a probability distribution over the topics conditional on a semantic unit. This distribution is given by

P(tj|uk) = P(tj)P(uk|tj) / sum (from j=1 to J) (P(P(tj)P(uk|tj))

In a similar way as discussed earlier, we use the dissimilarity between the distributions P(tj | uk) and P(tj) to measure the termhood of uk.

We first selected a number of OR journals. This was done based on the subject categories of Thomson Reuters. The OR field is covered by the category Operations Research & Management Science. Since we wanted to focus on the core of the field, we selected only a subset of the journals in this category. More specifically, a journal was selected if it belongs to the category Operations Research & Management Science and possibly also to the closely related category Management and if it does not belong to any other category. This yielded 15 journals, which are listed in the first column of Table 1.

In the first step of our methodology, the linguistic filter identified 2662 different noun phrases. In the second step, the unithood of these noun phrases was measured. 203 noun phrases turned out to have a rather low unithood and therefore could not be regarded as semantic units. ... The other 2459 noun phrases had a sufficiently high unithood to be regarded as semantic units.

In the third and final step of our methodology, the termhood of these semantic units was measured. To do so, each title-abstract pair in the corpus was treated as a separate corpus segment. For each combination of a semantic unit uk and a corpus segment si, it was determined whether uk occurs in si (nik = 1) or not (nik = 0). Topics were identified using PLSA. This required the choice of the number of topics J. Results for various numbers of topics were examined and compared. Based on our own knowledge of the OR field, we decided to work with J = 10 topics.

The evaluation of a methodology for automatic term identification is a difficult issue. There is no generally accepted standard for how evaluation should be done. We refer to Pazienza et al. (2005) for a discussion of the various problems.

We first perform an evaluation based on the well-known notions of precision and recall. We then perform a second evaluation by constructing a term map and asking experts to assess the quality of this map.

Precision is the number of correctly identified terms divided by the total number of identified terms.

Recall is the number of correctly identified terms divided by the total number of correct terms.

Unfortunately, because the total number of correct terms in the OR field is unknown, we could not calculate the true recall. This is a well-known problem in the context of automatic term identification (Pazienza et al. 2005).

To circumvent this problem, we defined recall in a slightly different way, namely as the number of correctly identified terms divided by the total number of correct terms within the set of all semantic units identified in the second step of our methodology. Recall calculated according to this definition provides an upper bound on the true recall. However, even using this definition of recall, the calculation of precision and recall remained problematic. The problem was that it is very time-consuming to manually determine which of the 2459 semantic units identified in the second step of our methodology are correct terms and which are not. We solved this problem by estimating precision and recall based on a random sample of 250 semantic units.

It is clear from the figure that our methodology outperforms the two simple alternatives. Except for very low and very high levels of recall, our methodology always has a considerably higher precision than the variant of our methodology that does not make use of PLSA.

A term map is a map, usually in two dimensions, that shows the relations between important terms in a scientific field. Terms are located in a term map in such a way that the proximity of two terms reflects their relatedness as closely as possible. That is, the smaller the distance between two terms, the stronger their relation. The aim of a term map usually is to visualize the structure of a scientific field.

It turned out that, out of the 2459 semantic units identified in the second step of our methodology, 831 had the highest possible termhood value. This means that, according to our methodology, 831 semantic units are associated exclusively with a single topic within the OR field. We decided to select these 831 semantic units as the terms to be included in the term map. This yielded a coverage of 97.0%, which means that 97.0% of the title-abstract pairs in the corpus contain at least one of the 831 terms to be included in the term map.

The term map of the OR field was constructed using a procedure similar to the one used in our earlier work (Van Eck and Waltman 2007b). This procedure relies on the association strength measure (Van Eck and Waltman 2009) to determine the relatedness of two terms, and it uses the VOS technique (Van Eck and Waltman 2007a) to determine the locations of terms in the map.

The most serious criticism on the results of the automatic term identification concerned the presence of a number of rather general terms in the map.

Another point of criticism concerned the underrepresentation of certain topics in the term map. There were three experts who raised this issue. One expert felt that the topic of supply chain management is underrepresented in the map. Another expert stated that he had expected the topic of transportation to be more visible. The third expert believed that the topics of combinatorial optimization, revenue management, and transportation are underrepresented.

As discussed earlier, when we were putting together the corpus, we wanted to focus on the core of the OR field and we therefore only included documents from a relatively small number of journals. This may for example explain why the topic of transportation is not clearly visible in the map.

When asked to divide the OR field into a number of smaller subfields, most experts indicated that there are two natural ways to make such a division. On the one hand, a division can be made based on the methodology that is being used, such as decision theory, game theory, mathematical programming, or stochastic modeling. On the other hand, a division can be made based on the area of application, such as inventory control, production planning, supply chain management, or transportation. There were two experts who noted that the term map seems to mix up both divisions of the OR field. According to these experts, one part of the map is based on the methodology-oriented division of the field, while the other part is based on the application-oriented division.

The experts pointed out that sometimes closely related terms are not located very close to each other in the map. One of the experts gave the terms inventory and inventory cost as an example of this problem. In many cases, a problem such as this is probably caused by the limited size of the corpus that was used to construct the map. In other cases, the problem may be due to the inherent limitations of a two-dimensional representation.

Our main contribution consists of a methodology for automatic identification of terms in a corpus of documents. Using this methodology, the process of selecting the terms to be included in a term map can be automated for a large part, thereby making the process less labor-intensive and less dependent on expert judgment. Because less expert judgment is required, the process of term selection also involves less subjectivity.

In general, we are quite satisfied with the results that we have obtained. The precision/recall results clearly indicate that our methodology outperformed two simple alternatives. In addition, the quality of the term map of the OR field constructed using our methodology was assessed quite positively by five experts in the field. However, the term map also revealed a shortcoming of our methodology, namely the incorrect identification of a number of general noun phrases as terms.

As scientific fields tend to overlap more and more and disciplinary boundaries become more and more blurred, finding an expert who has a good overview of an entire domain becomes more and more difficult. This poses serious difficulties for any bibliometric method that relies on expert knowledge.