顯示具有 co-word analysis 標籤的文章。 顯示所有文章
顯示具有 co-word analysis 標籤的文章。 顯示所有文章

2014年8月15日 星期五

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words. Journal of the American Society for Information Science and Technology, 62(10), 1933-1953.

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words.Journal of the American Society for Information Science and Technology,62(10), 1933-1953.

Scientometrics

圖書資訊學(LIS)為對於記錄下來的資訊(recorded information)和具有文化意義的文物與標本(culturally meaningful artifacts and specimens)有興趣的研究領域(Bates, 2010),包括的領域有檔案學(archival science)、 書目(bibliography)、文獻與文類理論(document and genre theory)、資訊學(informatics)、資訊系統(information systems)、知識管理(knowledge management)、圖書資訊學(LIS)、博物館研究(museum studies)、記錄管理(records management)和資訊的社會研究(social studies of information)。過去有許多研究嘗試定義與描述圖書資訊學的領域並且確認其中包含的研究主題,這些研究使用的方法相當廣泛,包含Järvelin & Vakkari (1990, 1993)採用內容分析(content analysis);Åström (2007, 2010)、Moya-Anegón, Herrero-Solana, & Jiménez-Contreras (2006)和 Persson (1994) 針對期刊或期刊文章進行書目計量分析 (bibliometric analysis) ; Moya-Anegón et al., (2006)和White & McCain (1998)針對作者進行書目計量分析 ;Åström (2002)、 Ding, Chowdhury, & Foo (2001) 和 Janssens, Leta, Glänzel, & De Moor (2006)利用從題名、摘要或全文抽取的詞語進行詞語的共現分析(co-word analysis) ;Sugimoto & McCain (2010)則是用索引詞語的三元共現分析(tri-occurrence analysis) ; van den Besselaar & Heimeriks (2006)利用詞語和參考文獻的組合進行分析;以及Sugimoto, Li, Russell, Finlay, & Ding, (2011)和 Sugimoto & McCain (2010)所使用的主題模型分析方法。

上述的這些方法,許多必須依賴於作者對於領域知識的了解,才能了解領域的主題與認知結構(cognitive structure),例如White & McCain (1998)基於最重要的作家的集群,觀察資訊科學由圍繞在一個微弱中心的許多專業所組成;Åström (2010)則是透過作者與期刊的映射圖說明這個領域的圖書館學(LS)和資訊科學(IS)之間具有差距。除了是認知結構較不直接的指標之外,引用分析另一個的問題是不同的次領域有不同的發表與引用實務。

論文題名包含許多能夠指出該文章內容的詞語(Buxton & Meadows, 1977; Meadows, 1998)。因此,本研究採用的方法是利用期刊論文題名上的重要詞語進行分析。分析的資料來自16種LIS期刊於1988到2007年發表的10344筆論文資料。

選取100個最常出現於題名的詞語。

本研究使用的分析技術包含詞語的相對頻率(relative frequency)並且根據詞語的共現進行叢集,最後並將詞語以及期刊與發表年度等進行多維尺度分析(multidimensional scaling, MDS),產生視覺化的結果。

詞語的共現分析以及階層式集群分析的結果發現三個主要分類LS(圖書館學)、IS(資訊科學)、SCI-BIB(科學計量學-書目計量學)以及兩個較小的分類資訊尋求行為(information-seeking behavior)和書目指導(bibliographic instruction)。LS可再細分為學術圖書館專業(academic librarianship)、公共圖書館專業(public librarianship) (包含館藏建立)、資訊素養和學校圖書館專業(information literacy and school librarianship, technology)、政策(policy)、全球資訊網(the web)、知識管理(knowledge management)、數位圖書館(digital libraries)、電子商務(e-commerce)、法律(law)以及學術出版(scholarly publishing)等主題。IS則包含資訊檢索(information retrieval)、網路搜尋(web search)、分類目錄(catalogs)以及資料庫(database)等主題。SCI-BIB也有書目計量指標(bibliometric indicators)、作者生產力(author productivity)與引用研究(citation study)等主題。整體的結構如下圖

從詞語的使用可以發現LIS中有某些持續出現的核心詞語,但也有一些詞語的使用在20年間有明顯的變化,這些都是與科技相關的(technologically related)詞語,這個現象符合Saracevic(1999)所宣稱的LIS是個科技驅動的(technology driven)領域。大致上來說,LIS內的改變可以從資料庫(database),到數位圖書館(digital libraries),到全球資訊網(the World Wide Web)等詞語使用的移轉上看得出來。

除了科技驅動的特徵外,LIS同時也有很大的範圍在討論資訊尋求行為,這是LS和IS都共同關心的課題。

A number of empirical studies of LIS have been conducted with the aim of describing and defining the field and identifying research areas within it. These studies applied a wide array of approaches: content analysis (Järvelin & Vakkari, 1990, 1993); bibliometric analysis of journals and journal articles (Åström, 2007, 2010; Moya-Anegón, Herrero-Solana, & Jiménez-Contreras, 2006; Persson, 1994); bibliometric analysis of authors (Moya-Anegón et al., 2006,White & McCain, 1998); co-word analysis of both index terms and words extracted from titles, abstracts, and full text (Åström, 2002; Ding, Chowdhury, & Foo, 2001; Janssens, Leta, Glänzel, & De Moor, 2006); tri-occurrence analysis of index terms (Sugimoto & McCain, 2010); analysis of word-reference combinations (van den Besselaar & Heimeriks, 2006); and topic analysis (Sugimoto, Li, Russell, Finlay, & Ding, 2011; Sugimoto & McCain, 2010).

Some notable studies of cognitive structure of LIS have interpreted topics post hoc, by assigning topicality based on knowledge of the author’s domain (e.g., White & McCain, 1998). In White and McCain’s influential visualization of LIS, they concluded that “information science lacks a strong central author, or group of authors, whose work orients the work of others across the board. The field consists of several specialties around a weak center” (p. 343). However, this analysis was based foremost on the clustering of authors, rather than topics. Similarly, Åström (2010) examined the divide between LS and IS components of the field by a bibliometric mapping of authors and journals. Topicality was assigned through expert knowledge of the domains in which these authors wrote and journals published.

Of the various components of textual documents, the titles, and the choice of words in them, are of particular importance. Title words function as “attention triggers” (Bazerman, 1985, 1988). They are devices for capturing interest in the world where information overload is a norm. Title words
have been called “signal-words”1 (Rip & Courtial, 1984) and “macro-actors” or “macro-terms”2 (Callon et al., 1983). Titles of journal articles themselves have undergone a change during the 20th century, becoming more informative, more specific, and containing a larger number of words that indicate article content (Buxton & Meadows, 1977; Meadows, 1998). Leydesdorff (1989) claims that “title words seem to offer a means of making visible the internal cognitive structure” (p. 217) of a discipline. He also claims that “word structure reflects internal intellectual organization in terms
of the codification of word usage in the relevant disciplines” (Leydesdorff, 1989, p. 221). 

Co-word analysis is based on co-occurrence of words (all words, or selected keywords) extracted from titles, abstracts, or text in general, or the index terms assigned by authors or indexers. Co-word analysis is a method that derives “higher level structures from word-occurrence patterns in text” (Chen, 2003, p. 139). Of particular importance in the context of this study is that co-word analysis is “a means to the elucidation of structures of ideas, problems, and so on, represented in appropriate sets of documents” (Whittaker, Courtial, & Law, 1989, p. 473). 

Although co-word analysis has its limitations, (e.g., Leydesdorff, 1997) primarily because of the
change of usage and meaning of words and the lack of context, such analysis has been considered particularly useful in tracking the development of scientific fields over time (Callon et al., 1991; Noyons & van Raan; Rip & Courtial, 1984), which represents another goal of this study.

Although citation analysis is not subject to the same limitation, it is a less direct indicator of cognitive structure. As already mentioned, studies using citations require post hoc assignment of topics. In addition, citation analysis of LIS is less effective in analyzing the cognitive structure of entire fields due to the different publication and citation practices of subfields, thus leaving even large subfields such as LS often invisible.

Selection of journals and articles. Articles from 16 LIS journals were chosen for inclusion in this study. The journals were selected from a ranked list of the most important journals in the field, according to deans and directors of American Library Association (ALA)-accredited, MLS programs in North America (Nisonger & Davis, 2005).

From this journal set, all research and review articles (10,344) published between 1988 and 2007 were included in the analysis.

Identification of the most frequently occurring LIS words and phrases. Word frequency is an important measure in content analysis. This measure is used to identify the most important research topics or concepts in a field by focusing on the most frequently occurring words.

In this study, we base all analyses on the 100 most frequently occurring LIS words or phrases. 

2014年1月26日 星期日

Glenisson, P., Glänzel, W., Janssens, F., & De Moor, B. (2005). Combining full text and bibliometric information in mapping scientific disciplines. Information Processing & Management, 41(6), 1548-1572.

Glenisson, P., Glänzel, W., Janssens, F., & De Moor, B. (2005). Combining full text and bibliometric information in mapping scientific disciplines. Information Processing & Management, 41(6), 1548-1572.

本研究以詞語共現分析(co-word analysis),將Scientometrics期刊2003年發表的論文,歸類為六個叢集。為了瞭解叢集結果的有效性,將這個結果與專家歸類的結果進行比較,同時也利用書目計量指標分析各個叢集。

Braam, Moed, and Van Raan (1991)建議利用詞語分析(word analysis)評估共被引叢集分析的結果,這些詞語利用書目紀錄裡的索引詞(indexing terms)和分類碼(classification codes)作為基礎。

在專家歸類方面,本研究援引Schoepflin and Glänzel(2001)研究的六個類別:數學模型與資訊計量學法則(Mathematical models/informetric laws)、個案研究(Case studies)、科學計量學的進展(Advances in Scientometrics)、指標工程(Indicator engineering)、社會學方法(Sociological approaches)與政策相關議題(Policy relevant issues)。加上近年興起的網路計量學(Webometrics)後,本研究用來進行專家的歸類的類別為:科學計量學的進展(Advances in Scientometrics)、實務論文與個案研究(Empirical papers/case studies)、數學模型(Mathematical models)、政策議題(Political issues)、社會學方法(Sociological approaches)以及資訊計量學與網路計量學(Informetrics/Webometrics)。下表是共詞分析與專家歸類的比較結果:

除了較大的A與E類別分布在多個叢集外,較小的類別大多集中內一個或兩個叢集上。

從六個叢集上的論文在專家以及它們的詞語網絡,可以將這些叢集分別:叢集1是書目計量學指標的方法學研究(methodological indicator research),這些指標用來測量發表活動(publication activity)以及引用影響(citation impact)的研究;叢集2大多為有關於國家和機構方面或科學領域的個案研究(case studies)與實務性的論文(empirical papers);叢集3和叢集1同樣是理論與方法學問題相關的論文,但更著重在資訊計量學法則(informetric laws)、頻率分布(frequency distributions)與多變化統計(multivariate stattistics)等先進方法學技術。叢集4是網路計量學和其他網路相關議題;叢集5是論文數較少的叢集,總共僅包括3篇論文,這些論文與共被引分析以及其他引用統計的分析有關;叢集6則是最大的叢集,包含的面向相當廣泛,從社會學、政策到科技等許多相關主題。從上述的分析,可以了解科學計量學目前主要的兩個面向是基於科學計量學標準技術的方法學研究和擴展傳統書目計量學範圍的實務研究。

接著利用平均參考文獻年齡(mean reference age)和連續出版品所占部分(share of serials)等書目計量學特徵分析上述的叢集結果。如下圖所示
在各個叢集裡,網路計量學具有低參考文獻年齡的特徵,並且連續出版品所占部分為中到高。政策議題相關的論文大部分具有相對低的連續出版品所占部分,但另有一群論文的連續出版品所占部分則明顯地高,因此相關的論文在圖形上分成兩個子叢集。至於科學計量學的先進方法與技術,除了少數例外,大部分的論文的平均參考文獻年齡在5到15年間,連續出版品所佔的部分則是在50%到90%間。實務性研究的論文在連續出版品所占部分的特徵分為兩群,一群的連續出版品所佔部分較低(<=55%),另一群則較高(>=67%),較低的一群與政策相關研究具有類似的特徵。

The question how bibliometric measures can, in turn, be assumed to reflect formal characteristics of documented scientific communication that might supplement results obtained from content-based analyses could also be answered in a positive way. Reference-based citation measures can help to fine-structure clusters determined on basis of co-word analysis.

Braam, Moed, and Van Raan (1991) suggested combining co-citation with word analysis in the context of evaluative bibliometrics to improve efficiency of co-citation clustering. The word analysis by Braam et al. used publication ‘‘word-profiles’’ that were based on indexing terms and classification codes.

Not much later, Noyons and Van Raan (1994) and Zitt and Bassecoulard (1994) demonstrated the appeal of plunging into contents by using keywords from both patent—and scientific literature to characterise the science-technology linkage.

The study by Schoepflin and Glänzel aimed at monitoring and characterising structural changes in the research profile in bibliometrics in the period 1980–1997. The authors created five categories, Mathematical models/informetric laws, Case studies, Advances in Scientometrics, Indicator engineering, Sociological approaches and Policy relevant issues. The term Webometrics did not yet appear in this scheme since at that time it was not yet established as a sub-discipline of scientometrics/informetrics.





We see classes S, M, I and P, admittedly all of smaller size, moderately to well conserved in the text-based cluster structure. Conversely, papers assigned to the larger classes A and E are heavily shifted around the text clusters.

The map in Fig. 6 represents the content structure of cluster 1 with altogether 9 papers. This cluster represents publications that are concerned with methodological questions related to bibliometric indicators. Indicator-related terms such as indicator names and terms relevant in the context of measuring publication activity and citation impact are close to the centre, and strongly interlinked. ... One could consider this cluster representing methodological indicator research.



Cluster 2 is dominated by empirical papers and case studies (cf. Table 3). ... The terms in this map are presented in Fig. 7 and relate above all to national and institutional aspects as well as to science fields. This is the cluster of case studies and traditional bibliometric applications.



Cluster 3 is a second theoretical/methodological cluster. Unlike the first one, this cluster relates to more advanced methodological techniques, such as informetric laws, frequency distributions and multivariate statistics. This cluster could be characterised as theoretical and mathematical issues in bibliometrics. The term structure is presented in Fig. 8.




Cluster 4 presented in Fig. 9 clearly represents webometrics and network-related issues. All terms are strongly interlinked. This cluster corresponds by and large to the category of Webometrics/Informetrics.



Cluster 5 with 3 papers is the smallest one. Co-citation analysis and the analysis of other citation statistics are the topic of these papers. The term structure (cf. Fig. 10) reflects the statistical vocabulary used in these studies. This cluster covers specific applications of statistical methods.



The last cluster with 30 papers (see Fig. 11) is by far the largest one. It comprises technology and innovation related studies, the science-technology interface and almost the complete Triple Helix issue can be found here (cf. Table 3). Also the sociological approaches are covered by this cluster. This cluster can be considered a borderland of classical scientometrics, namely the interdisciplinary approaches such as sociological, policy relevant and technology related issues.



The two large categories A and E covering 65% of all papers proved heterogeneous. Category A has (jointly with category M) three sub-clusters, namely, Cluster 1, 3 and 6, whereas Category E falls apart into three other sub-clusters: Cluster 2, 5 and 6. Policy relevant issues are also covered by clusters 2 and 6. Only Category I is represented by a corresponding co-word cluster, namely cluster 4.

The full text analysis substantiates that both methodological and empirical research have nowadays at least two different main focuses each, one is based on scientometric standard techniques such as classical indicators, the other ones are clearly broadening the scope of traditional bibliometrics.



As already seen in the pilot study, Webometrics is characterised by low reference age and medium–high share of serials (cf. Glenisson et al., 2005).

Most of the policy related issues are characterised by relatively low share of serials. Nevertheless, there is a group of papers with clearly higher share, too. This confirms the results of the full text analysis, namely that this category practically forms two sub-clusters.

The category Advances in Scientometrics proves strikingly homogeneous with several outliers only. Most of the A-class papers have, however, a mean reference age ranging between 5 and 15 years, with medium–high share of serials ranging between 50% and 90%.

The empirical groups proved heterogeneous, indeed. Regarding the share of serials this class forms two distinct sub-classes, particularly, one with low share (<=55%) and one with relatively high share (>=67%). The class with lower share has similar characteristics as the policy relevant class.

The question how bibliometric measures can, in turn, be assumed to reflect formal characteristics of documented scientific communication that might supplement results obtained from content-based analyses could also be answered in a positive way. Reference-based citation measures can help to fine-structure clusters determined on basis of co-word analysis.

2014年1月25日 星期六

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

本研究利用詞語共現分析(co-word analysis)技術,區分出六個圖書資訊學的研究主題:兩個書目計量學主題、一個資訊檢索主題、一個一般議題、一個網路計量學主題以及一個專利研究主題。

詞語共現分析根據詞語共同在文件出現的現象描述文件的內容,利用共同出現的相對強度呈現領域的概念網絡(concept networks)。目前已經有植物生物學(de Looze and Lemarie, 1997) 、凝態物理(Bhattacharya and Basu, 1998)、化學工程(Peters and van Raan, 1993)、資訊檢索(Ding, Chowdhury, and Foo, 2001)以及 醫學(Onyancha and Ocholla, 2005)等多個領域曾利用詞語共現分析技術來研究領域內的概念網絡。Van Raan and Tijssen (1993)討論基於詞語共現分析的書目計量在知識論的潛力(epistemological potentitals)。相較於共被引分析,詞語共現分析能應用在沒有引用索引的資料,而且共被引分析會因為在領域的變動與趨勢以及引用者的行為而變得複雜(Noyons & van Raan, 1998)。雖然Leydesdorff (1997)認為詞語的意義隨它們與其他詞語關係的頻率及其出現位置,會有所改變;但Courtial (1998)則是認為詞語共現分析中的詞語,並非做為用來代表某種意義的語言單位,而僅僅是文本間的連結指標。

本研究列舉幾個應用文字資訊為基礎的書目計量方法在圖書資訊學研究主題分析的研究:Courtial(1994)以詞語共現分析對這個領域進行探討,結果發現這個領域包含傳統圖書館學、資訊檢索、科學計量學、資訊計量學、專利分析以及最近興起的網路計量學。Glänzel及其同事整合全文為基礎的結構分析(full-text based structural analysis)和傳統的書目計量方法探討書目計量學及其次領域(Glenisson, Glänzel, and Persson, 2005; Glenisson, Glänzel, Janssens, and De Moor, 2005; Janssens, Glenisson, Glänzel, and De Moor, 2005)。

本研究所使用的分析技術包括:文本抽取(text extraction)、前處理(preprocessing)、多維度尺度(multidimensional scaling)以及Ward’s階層叢集(Ward's hierarchical clustering),並且利用向量空間模式(vector space model) (Salton & McGill, 1986)和隱藏語意分析(latent semantic analysis) (Deerwester et al., 1990)測量文件間相似程度的估計值。以論文彼此間的相似程度,將論文映射成二維圖形的結果如下,此圖形並且標示出每篇論文的期刊:

Scientometrics的論文主要分布在標示為1與2的兩個橢圓附近,橢圓1的主題為書目計量,橢圓2則為專利分析。橢圓5上的論文主要來自Information Processing and Management和Journal of the American Society for Information Science and Technology,其主題為資訊檢索。橢圓12的論文傾向於社會方面的主題,除了Journal of the American Society for Information Science and Technology以外,還包括Journal of Information Science和Journal of Documentation。正中央標示為14的橢圓,其主題與網路相關,所有的期刊均有這個主題的相關論文。

以Ward's叢集分析將所有論文進行歸類,最佳的結果共分為六個叢集。本研究並且根據每個叢集上論文的重要詞語以及中心的論文給予叢集的名稱。在二維圖形上標示各種叢集的結果如下:

六個叢集可以圖形上的斜線分為兩群,斜線以下為Bibliometrics1、Bibliometrics2和Patent Analysis,以上則為Webometrics、Information Retrieval和Social Aspects,但六個叢集中以Patent Analysis和其他叢集較分離。書目計量相關論文分為兩個叢集:Bibliometrics1和Bibliometrics2。Bibliometrics1與科學裡的合作關係(collaboration in science)、引用分析(citation analyses)和國家研究成效(national research performance)等主題相關,Bibliometrics2則主要為方法學和書目計量理論相關的論文。

為了找出各期刊分別著重的主題,除了比較上面的兩個圖形,另外還將叢集和期刊的關係映射成圖形。結果發現Information Processing and Management和Information Retrieval幾乎重疊,這個現象表示Information Processing and Management上的論文和Information Retrieval十分相關。Social Aspects和Webometrics相當靠近Journal of the American Society for Information Science and Technology、Journal of Information Science和Journal of Documentation三種期刊。事實上,除了Scientometrics以外,Social Aspects和其他期刊的距離大約相等。最後,Scientometrics則是落在Bibliometrics1、Bibliometrics2和Patent Analysis構成的三角形中心。

The optimum solution for clustering LIS is found for six clusters. The combination of different mapping techniques, applied to the full text of scientific publications, results in a characteristic tripod pattern. Besides two clusters in bibliometrics, one cluster in information retrieval and one containing general issues, webometrics and patent studies are identified as small but emerging clusters within LIS.

The method was developed by Callon, Courtial, Turner, and Brain (1983), more than two decades ago, for purposes of evaluating research. The methodological foundation of co-word analysis is the idea that the co-occurrence of words describes the contents of documents. By measuring the relative intensity of these co-occurrences, simplified representations of a field’s concept networks can be illustrated (Callon, Courtial, & Laville, 1991).

Van Raan and Tijssen (1993) have discussed the ‘‘epistemological’’ potentials of bibliometric mapping based on co-word analysis.

Leydesdorff (1997) analysed 18 full-text articles and sectional differences therein, and considered that the subsumption of similar words under keywords assumes stability in the meanings, but that words can change both in terms of frequencies of relations with other words, and in terms of positional meaning from one text to another. This fluidity was expected to destabilize representations of developments of the sciences on the basis of co-occurrences and co-absences of words.

However, Courtial (1998) replied that words, in co-word analysis, are not used as linguistic items to mean something, but as indicators of links between texts.

Many researchers have used this methodology to investigate concept networks in different fields, among others, de Looze and Lemarie (1997) in plant biology, Bhattacharya and Basu (1998) in condensed matter physics, Peters and van Raan (1993) in chemical engineering, Ding, Chowdhury, and Foo (2001) in information retrieval (IR) and Onyancha and Ocholla (2005) in medicine.

The reason why the emphasis has shifted from co-citation analysis to co-word techniques is twofold. The first reason is a practical one; co-word analysis allows application to non-citation indexes as well. The second relates to methodology; co-citation analysis complicates the combined analysis of field dynamics and trends in the actors’ activity (Noyons & van Raan, 1998).

Bonnevie (2003) has used primary bibliometric indicators to analyse the Journal of Information Science, while He and Spink (2002) compared the distribution of foreign authors in Journal of Documentation and Journal of the American Society for Information Science and Technology.

Bibliometric trends of the journal Scientometrics, another important journal of the field, have been examined by Schubert and Maczelka (1993), Wouters and Leydesdorff (1994), Schoepflin and Glänzel (2001), Schubert (2002), Dutt, Garg, and Bali (2003).

The main journals of the field were also analysed in terms of journal co-citation and keyword analyses (Marshakova, 2003; Marshakova-Shaikevich, 2005).

The co-citation network of highly cited authors active in the field of IR was studied by Ding, Chowdhury, and Foo (1999).

Finally, Persson (2000, 2001) analysed author co-citation networks on basis of documents published in the journal Scientometrics.

Courtial (1994) has studied the dynamics of the field by analysing the co-occurrence of words in titles and abstracts. Courtial described scientometrics as a hybrid field consisting of invisible colleges, conditioned by demands on the part of scientific research and end-users. Although this situation might have somewhat changed during the last decade, this conclusion illustrates how heterogeneous the much broader field of LIS – comprising subdisciplines such as traditional library science, IR, scientometrics, informetrics, patent analyses and most recently the emerging specialty of webometrics – nowadays is.

In recent papers, Glenisson, Gla¨nzel, and Persson (2005), Glenisson, Gla¨nzel, Janssens, and De Moor (2005), Janssens, Glenisson, Gla¨nzel, and De Moor (2005) have applied full-text based structural analysis in combination with ‘‘traditional’’ bibliometric methods to bibliometrics and its subdisciplines.

The full-text analysis consisted of text extraction, preprocessing, multidimensional scaling, and Ward’s hierarchical clustering (Jain & Dubes, 1988).

In short, the textual information is encoded in the vector space model using the TF-IDF weighting scheme, and similarities are calculated as the cosine of the angle between the vector representations of two items (see Salton & McGill, 1986; Baeza-Yates & Ribeiro-Neto, 1999).

The term-by-document matrix A is again transformed into a latent semantic index Ak (LSI), an approximation of A, but with rank k much lower than the term or document dimension of A. A latent semantic analysis is advisable, especially when dealing with full-text documents in which a lot of noise is observed.

One advantage of LSI is the fact that synonyms or different term combinations describing the same concept are mapped on the same factor, based on the common context in which they generally appear (Berry et al., 1995; Deerwester et al., 1990).

A lot of time was devoted to the detection of phrases. Since the best phrase candidates can be found in noun phrases, the programs LT POS and LT CHUNK4 have first been applied to detect all noun phrases in the complete document collection.

MDS represents all high-dimensional points (documents) in a two- or three-dimensional space in a way that the pairwise distances between points approximate the original high-dimensional distances as precisely as possible (see Mardia, Kent, & Bibby, 1979).

The agglomerative hierarchical cluster algorithm using Ward’s method (see Jain & Dubes, 1988) was chosen to subdivide the documents into clusters. ... One of the disadvantages of agglomerative hierarchical clustering is that wrong choices (merges) that are made by the algorithm in an early stage can never be repaired (Kaufman & Rousseeuw, 1990). What we sometimes observe when using hierarchical clustering is the forming of one very big cluster and a few small very specific clusters.

The journal Scientometrics can be largely separated from the other journals (which is also confirmed by the different term profile in the table of Appendix 1), and exhibits two different foci (best visible in Fig. 4).



The first ‘‘leg’’, indicated by the ellipse with number 1 and by and large containing the first focus of the journal Scientometrics, clearly contains papers in bibliometrics. The 10 best TF-IDF terms for ‘‘leg’’ #1 are: citat, cite, impact factor, self citat, co citat, scienc citat index, citat rate, isi, countri and bibliometr.

The second ‘‘leg of Scientometrics’’, indicated by number 2, is characterised by the best terms patent, industri, biotechnolog, inventor, invent, compani, firm, thin film, brazilian and citat. The JIS paper (#3) embedded in this patent ‘‘leg’’ might be considered an outlier for that journal, but it was put in the right place since it is concerned with ‘‘The many applications of patent analysis’’ (Appendix 2: Breitzman & Mogee, 2002).

An important focus of LIS is indicated by ellipse #5 and can be profiled as ‘‘Information Retrieval’’ (IR) when looking at the highest scoring terms: queri, search engin, web, node, music, imag, xml, vector and weight.

The fourth distinguishable subpart of LIS (#12) is about digit, internet, servic, seek, behaviour, health, knowledg manag, organiz, social and respond; so encompassing the more social aspects.

The remaining large subpart is somewhat the central part (#14). It consists of papers leading to a mean profile containing the terms web, web site, classif, domain, web page, languag, scientist, region, catalog, and web impact factor.

The term network of Cluster 1 allowed the conclusion that the papers belonging to this cluster are concerned with domain studies, studies of collaboration in science, citation analyses, national research performance and similar issues.



The medoid is a paper by Persson et al. on ‘‘Inflationary bibliometric values: The role of scientific collaboration and the need for relative indicators in evaluative studies’’ (Appendix 2: Persson et al., 2004). This is a methodological paper with strong implications for research evaluation, combining research collaboration with citation analysis and construction of national science indicators.

The smaller bibliometrics cluster (Cluster 3: manually labelled as ‘‘Bibliometrics2’’) is of more methodological/theoretical nature.




The medoid is the state-of-the-art report ‘‘Journal impact measures in bibliometric research’’ (Appendix 2: Gla¨nzel & Moed, 2002).

The term networks for the two bibliometrics clusters just described contain a few overlapping terms (bibliometr, chemistri, citat, citat rate, cite, cluster, countri, impact factor, isi, physic, rank and scienc citat index). The MDS plot of Fig. 15 confirms that there is no clear border between Bibliometrics1 and Bibliometrics2, but that there is a gradual transition.

The almost tiny Cluster 2 (19 papers, Fig. 10) represents patent analysis.


A paper on ‘‘Methods for using patents in cross-country comparisons’’ forms the medoid of this cluster (Appendix 2: Archambault, 2002).

Cluster 4, with 282 papers, is the largest one. We have labelled it ‘‘Information Retrieval’’.


The medoid paper is entitled ‘‘Querying and ranking XML documents’’ (Appendix 2: Schlieder & Meuss, 2002).

Cluster 5, with 62 papers, belongs to the small clusters. Both terms and papers close to the medoid characterise this cluster as ‘‘Webometrics’’.


The medoid paper is entitled ‘‘Motivations for academic web site interlinking: evidence for the Web as a novel source of information on informal scholarly communication’’ (Appendix 2: Wilkinson et al., 2003).

Cluster 6 (213 papers) proved to be the most heterogeneous cluster. We have labelled it ‘‘Social’’, however, we could also have called it ‘‘General & miscellaneous issues’’.



‘‘Approaches to user-based studies in information seeking and retrieval: a Sheffield perspective’’ is the title of the medoid paper (Appendix 2: Beaulieu, 2003).


The Patent cluster can be clearly separated from the rest of LIS. The subspace under the line is almost completely occupied by Bilbiometrics1, Bibliometrics2 and Patent.




IR and IPM almost collide in this 2D projection (Fig. 20). This means that Cluster 4 (‘‘IR’’) is very close to the scope of this journal.

The ‘‘Social’’ cluster with general and miscellaneous topics as well as ‘‘Webometrics’’ are close to JIS, JDoc and JASIST, too. Moreover, the ‘‘Social’’ cluster is almost equidistant to all traditional journals in Information Science.

The remaining three clusters, namely Bibliometrics1, Bibliometrics2 and Patent, form a triangle in the centre of which the journal Scientometrics is located. The relatively large distances among these clusters and between each cluster and the journal, strongly indicate that a quite large spectrum of bibliometric, technometric and informetric research using different vocabularies is covered by the journal Scientometrics. This observation is in line with the findings by Schoepflin and Gla¨nzel (2001) that scientometrics consists of several subdisciplines such as informetric theory, empirical studies, indicator engineering, methodological studies, sociological approach and science policy; and that case studies and methodology became dominant by the late 1990s. At the end of the 1990s, also technology related studies based on patent statistics became an emerging subdiscipline of the field.

We have found two clusters in bibliometrics, of which a big one in applied bibliometrics/research evaluation and a smaller one in methodological/theoretical issues; also we have found two large clusters in information retrieval and general and miscellaneous issues and, finally, two small emerging clusters in webometrics and patent and technology studies. Within the IR cluster, we have found a small subcluster on music retrieval, which might be a temporary phenomenon since the journal JASIST has published a special issue on this topic.

According to the expectation, IR, General issues and Webometrics were represented by four of the five journals, namely JIS, IPM, JASIST and JDoc, while the two bibliometrics and the patent clusters were the domain of the journal Scientometrics.

2013年12月19日 星期四

Zhu, D. and Porter, A. L. (2002). Automated extraction and visualization of information for technological intelligence and forecasting. Technological Forecasting & Social Change, 69, 495-506.

Zhu, D. and Porter, A. L. (2002). Automated extraction and visualization of information for technological intelligence and forecasting. Technological Forecasting & Social Change, 69, 495-506.

information visualization
本論文認為因為需要處理大量的文字資料、處理時要能快速以及呈現結果時需要生動並且能夠理解等三種需要,建議利用文字探勘(text mining)技術以及書目計量指標(bibliometric indicators)分析大量的文字資料庫,產生一系列的技術地圖(technology maps)和創新指標(innovation indicators)。其中產生各種技術地圖的技術包括在圖形上將資料項目對映到適合位置的MDS技術以及連結相關項目對映節點的路徑消除(path-erasing)演算法,並且本論文也建議使用詞語的共現資訊做為資料項目間相關程度的評估參考。
Three factors could enhance managerial utilization: capability to exploit huge volumes of available information, ways to do so very quickly, and informative representations that help manage emerging technologies.
Empirical analysis of emerging technologies poses a number of challenges to analysts. In particular, we note the need to:
1. digest enormous amounts of available information,
2. do so rapidly,
3. present findings vividly and understandably.
A third hard-earned lesson gained from our developmental experiences with ‘‘bibliometrics’’ (counting bibliographic activity) and ‘‘text mining’’ has been that TF-related results must be easily understood and must directly relate to a user’s perceived information needs.
This paper reports on efforts to address these three factors via partially automated processes to generate helpful knowledge from text quickly and graphically. We first illustrate a process to generate a family of technology maps that help convey emphases, players, and patterns in the development of a target technology. Second, we exemplify the generation of particular ‘‘innovation indicators’’ that measure particular facets of R&D activity to relate these to technological maturation, contextual influences, and market potential.
In sum, then, we seek to respond to these challenges—analyzing large text resources, rapidly, to generate compelling findings—to enhance TF (including competitive technological intelligence, technology foresight, etc.). Our approach, called technology opportunities analysis (TOA), seeks to facilitate this process by profiling search sets of bibliographic abstracts on technologies of interest.
The TOA process entails these main steps:
1. Search and retrieve text information, typically from large abstract databases.
2. Profile the resulting search set. VantagePoint applies a combination of machine learning, statistics, and natural language processing to yield what van Raan (1992) call a mix of ‘‘one-dimensional’’ descriptions (lists) and ‘‘two-dimensional’’ relationships (matrices). Profiling may focus on documents. Or, it may focus on concepts (e.g., principal components analysis (PCA) to group related terms as conceptual clusters). A third choice is a combination—seeking to link documents to concepts.
3. Extract latent relationships. VantagePoint applies iterative principal components analyses to uncover links among terms and underlying concepts.
4. Represent relationships graphically. Generation of ‘‘mapping’’ and ‘‘indicators’’ are elaborated in the following sections.
5. Interpret the prospects for successful technological development. This typically entails integrating the bibliographic search set analyses with expert domain knowledge (interviews).
We have developed a partly automated process to do so based on ‘‘co-occurrence’’ information. Co-occurrence is based on the pattern of terms occurring together in the records. If two terms occur together in the records more frequently than expected, there is a presumption of relationship between them. Terms can include authorship (also organizational affiliation, nationality) or ‘‘keywords’’ (subject index terms), or noun phrases generated from titles or abstracts using our natural language processing (NLP) routine (cf., Refs. [18,20]).
Effective visualization of the basic co-occurrence and correlation matrix information entails a sequence of analyses:
1) a new two-step multidimensional scaling (MDS) algorithm,
2) an improved path-erasing algorithm,
3) a routine to determine and display size (relative frequency of occurrence),
4) macros to create maps in VantagePoint, Microsoft Word or MS PowerPoint,
5) a routine to consolidate duplicate principal components (in the mapping process),
6) an algorithm to automatically name principal components,
7) an algorithm to cut off principal components to just include high-loading terms (the last three steps are needed for principal components maps; cf., Refs. [16,18]),
Our routine generates various maps, such as:
1. principal components map [represents the relationships among conceptual clusters];
2. keywords map [represents the relationships among frequently occurring subject index terms, title phrases, or whatever terms are chosen];
3. affiliations map [represents the relationships of affiliations’ research topics, based on terms they use in their documents—see Fig. 1];
4. authors map [analogous to affiliations map, but for individual researchers];
5. countries map [analogous to affiliations map];
6. sources (e.g., journals) map [analogous to affiliations map].
Fig. 1 shows an affiliations (organizations) map for the ‘‘Nanotechnology’’ topic. Displayed are the most prolific publishers abstracted in INSPEC for 1998. Along with the organizational name are shown the three keywords most frequently used in its publications in the search set. The size of a node reflects the number of publications. Positioning is determined using our MDS and path-erasing algorithm.
In essence, the challenge is to reduce n-dimensional (in this case, n equates to 40-dimensional since there are some 40 affiliations’ similarity being represented) to 2-D or 3-D. MDS is the generally favored approach to accomplish this. In MDS, an important parameter called stress is used to control its procedures. The process of generating a MDS map seeks the optimum location for each element in the map by minimizing the stress. ... We have devised a ‘‘step-by-step’’ search algorithm. This algorithm is effective at finding the global stress minimum, although it usually consumes more CPU time than the ‘‘steepest descent’’ algorithm.
Therefore, we have added an additional representational element, connecting links, based on a ‘‘path-erasing’’ algorithm. This is built on a proximity matrix among the elements. Its logic is as follows:
1. connect all elements in the proximity matrix together,
2. set a series of thresholds to erase the connecting lines one by one,
3. devise a suitable stop criterion.
The partially automated processes presented provide ‘‘value-added’’ knowledge from bibliographic text mining. The family of maps allows a user to gain an intuitive feel for R&D activity.
We suggest that development of routines to generate particular representations—technology maps and innovation indicators—automatically can enhance the applicability of text mining and bibliometrics to TF. ... However, scripting the production of these visualizations can facilitate provision of empirically based, vivid TF findings, in a timely manner, to inform decision making. That could dramatically increase the utilization of TF in management of technology

2013年12月2日 星期一

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

科學映射圖(scientific mapping)能夠科學結構(scientific structure)視覺化,幫助使用者確認科學主題(scientific themes)並從而發現新知識的有用工具之一。過去的研究曾經使用過作者、文章與等映射單位。在計算映射單位之間的關連,Börner, Chen, and Boyack (2005) 將關連性的測量方法(relatedness measures)分為引用連結(citation linkages)與共現相似性(co-occurrence similarities)等兩大類,而本研究則將目前常用來評估作者間的關連分為直接引用(direct citation)、共被引分析(cocitation analysis)、合著分析(co-authorship analysis)、書目耦合分析(bibliographic coupling analysis)以及共詞分析(co-word analysis)等五種方法。也有研究以發展出整合文字內容與連結的測量方法來計算期刊(Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005)與文章(Liu et al., 2010)間的關連。本研究建議兩種以詞語為基礎並利用向量空間模式(vector space modeling)的方法和另一種基於LDA(latent Dirichlet allocation)的主題模型方法來測量作者之間的關連。本研究將第一種方法稱為靜態(static)的特徵,以每位作者曾寫過的論文內容為基礎產生代表這位作者的特徵向量,也就是代表這位作者的特徵向量是所有他寫過的論文的特徵向量總和,任何兩位作者之間的關連是對應於他們的作者特徵向量之間夾角的餘弦值(cosine value)。第二種方法則是動態(dynamic)的特徵,如果兩位作者之間沒有合著的論文,他們之間的關連仍然是他們的作者特徵向量之間夾角的餘弦值,但如果他們曾經合著過,在計算他們之間的關連時,先將他們合著論文的特徵向量排除在他們的作者特徵向量之外,在進行餘弦值計算,所以在計算每位作者和其他作者之間關連時所使用的作者特徵向量可能是變動的,因此稱為動態。基礎的主題模型假設每一個論文都是主題的混合(mixture),而每一個主題則都是詞語的混合。對於每一個論文,它的主題混合由一個已知參數α的Dirichlet分布所產生;每一個主題的詞語混合則由另一個已知參數β的Dirichlet分布所產生。在產生論文d前先根據Dirichlet分布Dir(α)取樣產生它的主題混合θd,然後再產生這個論文裡的每一個詞語,每一個詞語的產生是根據從主題混合θd中取樣得到的主題z以及其相對應的詞語混合ϕk所產生。本研究採用Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010)將作者資訊加入而擴充的LDA模型-- 作者-主題模型(author-topic model),這個模型假定每個作者是由一個已知參數α的Dirichlet分布所產生的主題混合。假設一個論文的作者群為ad ,在產生這個論文的每一個詞語時,首先從ad 中隨機抽取一個作者x以及他的主題混合θx,然後其主題z便由θx取樣產生。本研究利用Gibbs取樣(Gibbs sampling, Griffiths & Steyvers, 2004)進行作者-主題模型推論,產生包含每一個主題在詞語上的分布情形以及對每一位作者產生他在各主題上的分布情形等結果。因此利用作者-主題模型可以根據他們在主題分布的相似度測量他們的關連。

本研究的資料範圍為2000到2010年出版的圖書資訊學相關的八種主要期刊的 5227筆書目紀錄,從其中的 6282位不同的作者內選取50位最多產的作者。利用靜態特徵、動態特徵、主題模型和共被引分析等四種方法測量多產作者之間的關連並利用MDS (multidimensional scaling)和階層式叢集分析(hierarchical cluster analysis)進行視覺化。本研究在利用主題模型測量作者之間的關連時使用以下的參數,α設為50/K,其中的K是主題的數量,本研究設為20,β設為0.01,Gibbs取樣的迭代(iteration)次數設為1000次。針對每一對作者的四種關連測量方法所得到的值進行相關分析(correlation analysis),結果發現靜態特徵與動態特徵之間有最高的相關值,主題模型和其他兩種以內容為基礎的測量方法的相關值也較共被引方法來得高。四種測量方法皆可以發現LIS領域的兩大主軸:一個主軸是資訊檢索(information retrieval)與網路研究(web studies),另一則是科學評鑑(scientific evaluation)的測量指標(metrics)研究,LDA模型則在階層式叢集分析上有最連貫的結果。另外,以內容為基礎的方法比以引用為基礎的方法更容易解釋產生的結果。

In this study we present static and dynamic word-based approaches using vector space modeling, as well as a topic-based approach based on latent Dirichlet allocation for mapping author research relatedness.

Outcomes for the two word-based approaches and a topic-based approach for 50 prolific authors in library and information science are compared with more traditional author cocitation analysis using multidimensional scaling and hierarchical cluster analysis.

Science mapping is one of the most useful tools to visualize scientific structure. It helps to identify scientific themes, and discover new knowledge.

The unit of interest for mapping may include authors, articles, and journals.

To date, five approaches have been used to measure the relatedness between authors, where the nature of the relationship studied is based on the data used: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling analysis, and co-word analysis.

Recently, more sophisticated hybrid methods (i.e., using textual content and citations) have been applied to the mapping of articles (Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005) and journals (Liu et al., 2010).

As an initial investigation of these topics, our focus will be on authors whose publications appear in the highest impact library and information science journals.

In reviewing visualization studies for knowledge domains, Börner, Chen, and Boyack (2005) categorized relatedness measures into two broad categories: citation linkages and co-occurrence similarities.Within the relatedness measures, five basic approaches were identified: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling, and co-word analysis.

Direct citation accounts for the relatedness between a citing work and a cited work based on citing behavior. ... Shibata, Kajikawa, Takeda, and Matsushima (2008) explored citation networks for two research domains and divided the networks into clusters in order to identify research fronts. Direct citation has not attracted wide attention. One possible reason may be its requirement for a very long time window to obtain a sufficient linking signal for clustering (Boyack & Klavans, 2010).

The idea that two articles that share the same references are related, referred to as bibliographic coupling, was outlined by Kessler (1963). The more references two articles have in common, the more closely related they are thought to be. Note that this list is static over time because references within articles do not change. With the interrelation of this link, scientific products can be ordered into groups. Weinberg (1974) reviewed the theory and practical applications of bibliographic coupling and granted the usefulness of the method. More recently, Zhao and Strotmann (2008) aggregated bibliographic coupling at an author’s oeuvre (body of work) level, which they called author bibliographic-coupling analysis (ABCA). They found ABCA can provide an effective picture of current active research in a field.

Cocitation analysis, introduced by Small (1973), is probably the most influential approach for assessing relatedness measures. If two articles are cited by the same third article, these two articles are co-cited. The assumption is that the appearance of two articles in the same reference list indicates a semantic association between the articles. Unlike traditional bibliographic coupling, cocitation is a dynamic relationship based on the citing authors. New citing authors can change the cocitation relationship. This feature is important because science is developing continuously. Relationships among scientific units being studied should be able to incorporate this dynamic change.

White and Griffith (1981) first applied cocitation techniques to authors, called author cocitation analysis or ACA. The essential transformation is to consider “Author” as a body of writings by a person (i.e., an oeuvre). So the cocitation of authors applies to any work by any author being co-cited with any work by another author.

Since then, a number of studies have been conducted using variations of the ACA method, including normalization (Ahlgren, Jarneving,&Rousseau, 2003; Leydesdorff&Vaughan, 2006; White, 2003; van Eck & Waltman, 2009), author counts (Zhao & Strotmann, 2011), and last-author ACA (Zhao & Strotmann, 2010).

One disadvantage of cocitation analysis is the lack of cognitive interpretation of the relatedness of the co-cited units. Without enough domain knowledge, one can hardly interpret the cocitation map.

Leydesdorff (1987) argued that cocitation maps only partially represent the structure of science.

A co-authorship relationship is established when authors co-publish a paper. Glänzel (2001) studied international co-authorship links to reveal the structures in international collaborations. Liu, Bollen, Nelson, and Van de Sompel (2005) constructed a network with co-authorship relations in the field of digital libraries. Ding (2011b) studied scientific collaborations and citation patterns of researchers and combined the results with a topic model approach to examine collaborations among researchers who share similar and different research interests.

It is this feature of co-authorship that makes co-authorship analysis more revealing of a social network rather than a scientific structure.

Co-word analysis collects evidence of relatedness from co-occurring keywords from different articles. Compared with the approaches introduced earlier, co-word analysis directly uses actual contents to measure relatedness, whereas the others find indirect evidence through citation and co-author relations. An obvious advantage of co-word analysis is that relatedness can be interpreted directly according to document contents.

Coulter, Monarch, and Konda (1998) mapped the discipline of software engineering with co-word analysis. Indexing terms from the ACM Computing Classification System were used as the unit of analysis. Ding, Chowdhury, and Foo (2001) conducted a co-word analysis on a sample of 2,012 articles from the Web of Science (WoS) to reveal themes of information retrieval research.

Leydesdorff (1997) noted that the meaning of words change from position to position and from one text to another. He also suggested this change will destabilize the science map produced by co-word analysis.

Another disadvantage of using indexer-assigned keywords as the source for co-word analysis is the “indexer effect” (Law & Whittaker, 1992), which creates bias through factors such as the artificiality of an indexing language, delays in changes to the indexing language to reflect the current state of a discipline, and subjectivity in the assignment of index terms.

In the vector space, a number of documents constitute a document space. The centroid of the document space is a summarization of the characteristics of the space. It represents the average vector for a group of documents.

Each author will be viewed as a document space consisting of the articles he/she has written. This space is a subspace of the collection space, named the author space. The centroid of the author space will be used to represent the author. The relatedness between authors will be measured through the similarity between the centroids of their author spaces.

The topic model is an improvement over the basic vector space model in terms of relieving the independence assumption and capturing the term associations. Instead of assuming independence among terms, the topic model assumes exchangeability among terms in documents, which is a much looser assumption.

Early works on the topic model include latent semantic indexing (LSI) by Deerwester et al. (1990) and the probabilistic LSI (pLSI) by Hofmann (1999). LDA is a more recent technique proposed by Blei, Ng, and Jordan (2003). It has an advantage over LSI in explicitly modeling the latent topics, and over pLSI in solving the overfitting problem (i.e., a model with too many parameters).

The LDA model treats a document as a mixture of topics and a topic as a mixture of terms. Each document (i.e., a mixture of topics θ) is generated from a latent Dirichlet distribution with a prior of α, and each topic (i.e., a mixture of terms ϕk) is generated from a Dirichlet distribution with a prior of β. The generation process entails, first, sampling a document θd from Dir(α). At each position of a word in a document, a topic z is selected according to θd, and a word w is selected according to z and ϕk.



Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010) extended the original LDA model to include authors and proposed the author-topic model (Figure 2). This model includes authorship information in the generative process. Each document has a number of authors ad. Each author is considered as a distribution of topics drawn from a Dirichlet distribution with a prior of α. For each word in a document, an author x is randomly drawn from ad and the topic distribution associated with this author is θx. Then a topic z is selected the same way as in a LDA model to generate the observed word w.



The advantage of this author-topic model is that it adds authorship information to the model, so that the topics are learned and assigned to documents accordingly. In the output of this model, each author is a distribution of different topics; each topic is a distribution of terms. As the purpose of the current study is to measure the relatedness of authors, the author-topic model will be appropriate to produce author similarities based on their topics.

Gibbs sampling (Griffiths & Steyvers, 2004) is used to estimate the parameters in the model.

Table 1 lists the eight journals selected for inclusion in the study. ... Bibliographic records for documents published in these journals between 2000 and 2010 were downloaded. Records downloaded were further limited to three document types: articles, proceedings papers, and reviews. ... In total, 5,227 records were downloaded from WoS. The raw WoS records were processed, and only three fields were kept: the article title (i.e., “TI” field), the Keywords Plus (i.e., “ID” field), and the abstract (i.e., “AB” field). The records then were indexed with the widely used Lemur information retrieval toolkit (http://www.lemurproject.org/). Stop words were removed and stemming was applied.



From the 5,227 records downloaded, we were able to identify 6,282 different author names using string matching. Because it is impractical to map all of the authors in our collection, we selected the 50 most prolific authors according to the WoS “analyze results” function. ... We selected the most prolific authors because the more an author writes, the better the algorithm used “understands” her/his interests, and thus the more accurate our assessment will be.

For each author in our author list we then generated an author space consisting of all the articles he/she wrote. TF*IDF term weighting was employed to assign term significance in the space. Terms that were single characters or only consisted of digits (e.g., “2001”) were filtered out. We believe that these terms add noise to the space rather than meaning. The relatedness between authors is measured through the cosine between the centroids of the author spaces.

One could argue that this creates a biased assessment of the strength of the relationship because there is an exact match for the text of the co-authored publications that creates a stronger bond than for two authors who have published in a common area but did not collaborate. On the other hand, the simple fact that the collaboration has resulted in one or more co-authored documents should be acknowledged as a strong tie between the authors.

In a static space, each author has her/his own space that consists of her/his articles. This space does not change when measuring author relatedness. ... The relatedness of authors will include the similarity arising from the strength of the co-authorships.

Conversely, in the dynamic author space, the author spaces depend on a pair of authors. Co-authored articles by the pair of authors are excluded. In this case, each author may have a different author space when measured with different authors.

The vector space model provides a number of readily available measures of relatedness. The most popular is the cosine measure, which measures the cosine of the angle formed by two vectors in the space. It basically measures the term weight distribution between two vectors. The more similar the distribution is, the higher the cosine value is expected to be.

Gibbs sampling (Griffiths & Steyvers, 2004) was used to estimate the parameters in the author-topic model. We set the number of iterations to 1,000. The hyperparameter α was set to 50/K where K is the number of topics and hyper β is set to 0.01.We tested different K, or number of topics, values and decided to report the results from K = 20 because it produced the most reasonable outcome by our judgment.

The topic model toolbox was employed to perform the learning process (http://psiexp.ss.uci.edu/research/programs_data/toolbox.htm).

An author-topic LDA model (Rosen-Zvi et al., 2010) was trained on our collection and a pair-wise cosine similarity measure comparison of the 50 authors was conducted, resulting in a symmetric matrix of similarity values based on the LDA modeling. Similarity matrices were also calculated for both the static and dynamic author spaces. Multidimensional scaling was used to visualize the relationships among the authors. ... Because the data represent a type of similarity measure, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006). To provide additional insights into the grouping of the authors, hierarchical cluster analysis (complete linkage method) was used in SPSS to superimpose groups of authors on the MDS maps to provide an additional means to assess the coherence in the resulting proximities between authors.

After tokenization of the field contents, 916,383 tokens, or individual words, were identified; the number of unique tokens, or distinct words, was 12,537. The average document length was 175.32 tokens.

An examination of the pair-wise correlation of these author relatedness measures reveals significant and moderate level correlations between the word-based, topic-based, and author cocitation measures (Table 5). It is not surprising that the static author map has a high correlation with the dynamic author map (Kendall’s tau b = 0.971). Similarly, the correlations among the three content-based approaches are generally higher than their correlations with the cocitation approach. This provides preliminary evidence that they measure different types of relationships.

In all cases, the largest singular group consists of authors who work with different aspects of metrics-based studies, which is labeled as “Informetrics” in general in the two word-based maps and “Scientific impact evaluation” in the other two maps. This labeling indicates that the metrics-related topics have been a frequently investigated theme by the prolific authors in the selected journals during the first decade of the 21st century.

It is also noteworthy that the topic groupings of each of the maps largely aligns along the horizontal or vertical axis, with one side representing information retrieval (system and behavior) and web studies, with the other side corresponding to metrics-based or scientific evaluation studies.

As is shown from the maps, the static map (Figure 3) and dynamic map (Figure 4) are generally consistent in terms of the location of the authors, which indicates that the exclusion of similarities resulting from collaborations does not affect the overall layout. However, drastic changes may happen to individuals who have collaborated frequently with another author.

At the four-cluster agglomeration, the LDA map (Figure 6) provides the most coherent representation of the author map in relation to the generated clusters. At the two-cluster agglomeration, the clusters are neatly divided along the vertical axis, with metrics-related research represented on the left, and web and information retrieval-related themes on the right. Although the group membership of some individuals is still debatable, such as “Ingwersen_P” in the “Scientific impact evaluation” group given that he has also published in information retrieval and webometrics, the overall layout of the LDA map does provide semantically meaningful relationships.

Of the five author relatedness methods discussed earlier, only co-authorship provides a direct connection between authors.

Cocitations are contributed by third parties.

Direct citations reflect an author’s assessment of relatedness to a cited author or work but are still based on perception or the subjectivity inherent in citer motivation (Bornmann & Daniel, 2008).

This is also the case for bibliographic coupling, where the strength of the relationship is assessed by the overlap of references selected by two authors.

Co-word or topic-based studies can be argued to be the least influenced by citing behavior because they rely solely on the words developed by the authors themselves.

The newly proposed content-based approaches overcome several limitations of the more traditional cocitation approach.

In addition to avoiding citer subjectivity inherent in citation-based data, the links between authors will be more interpretable compared with the cocitation maps. The top terms/topics will be identifiable to help interpret the links between authors.

The content-based methods do not require an author to be cited in order to be included in the map. As long as the author has some publication record, her/his relatedness with other authors can be identified. This provides the opportunity for researchers who have not been widely cited to be included in the author map.

Furthermore, cocitation analysis outcomes may be affected by limited numbers of citations that do not reflect the true strength of the relationship between authors. This can be seen when comparing the cocitation outcomes with the topic-based outcomes, where several authors with low citation counts, and therefore low cocitation counts, end up at the periphery of the map. For the LDA outcome, these authors are more centrally situated among authors with similar topic areas.

The word-based and topic-based methods can be considered an extension of co-word analysis, where words are used to determine the relatedness of authors.

In Healey, Rothman, and Hoch (1986), a paradox is introduced: if a map represents a field that is already known to experts, then it is useless because it does not reveal anything new; if the map deviates from the expectation of the experts, then its outcome is questionable.

This initial investigation, which compares prolific authors from LIS, demonstrates: (1) the potential for more topically meaningful outcomes from the new methods when compared to more traditional cocitation analysis; (2) the topicbased method using LDA for the data used in this study produces more distinctive clusters and reasonable results than the two word-based approaches.

2013年4月8日 星期一

Rorissa, A. and Yuan, X. (2012). Visualizing and mapping the intellectual structure of information retrieval. Information Processing and Management, 48, 120-135.

Rorissa, A. and Yuan, X. (2012). Visualizing and mapping the intellectual structure of information retrieval. Information Processing and Management, 48, 120-135.

network analysis

Chen (2006)說明一個學術領域或學科的知識基礎(intellectual base)與研究前沿(research front)的區別:研究前沿是一個專業(specialty)當前最先進的技術狀態(state of the art),由研究前沿引用所構成的部分則是它的知識基礎。通常在分析學術領域或學科的知識基礎時,大多利用期刊的引用資料,並且使用群集分析(cluster analysis)、多維分析(multidimensional analysis)與其他技術將引用資料的視覺化 (例如:Chen & Kuljis, 2003; Chen et al., 2010; Ding et al., 2000; McKechnie, Goodall, LajoiePaquette, & Julien, 2005; Tang, 2004; White & Griffith, 1981; White & McCain, 1998)。White (2003)則是使用尋徑網路(Pathfinder Networks)技術將White & McCain (1998)的資料繪製成圖書資訊學領域的科學映射圖。這些技術也都被應用於軟體工具的製作並且提供研究人員自由使用,例如CiteSpace。將引用資料等資訊繪製成科學映射圖的研究興趣提升可以歸納以下的原因:可使用的引用資料來源以及其他廣泛出現;許多提供視覺化與映射圖的電腦應用程式可以自由使用;對不斷增加數量的電子資料需要具有包容性與容易使用的管理與理解方法。

這篇研究利用2000到2009年間的資訊檢索領域的引用資料進行一系列的分析與視覺化,所利用的資料來源是Web of Science,共使用48,390筆書目紀錄,使用的視覺化工具是由Chen(2004a; 2004b; 2006)提出的CiteSpace (http://cluster.cis.drexel.edu/cchen/citespace/)。研究的項目與結果如下:

1) 作者合著網路(co-authorship network): 以32位發表10篇以上的較高生產力作者為分析對象,,其中一半來自於電腦科學,另一半則為資訊科學。只要兩位作者曾一起出現在一篇論文中便在他們的映射點間建立連結。網路建立好之後,測量每位作者映射點的中介性,找出中心作者。結果可以發現兩個作者數目超過八位的較大作者叢集,中心分別是 Jarvelin K和Chen HC,前者來自資訊科學,而後者則是來自電腦科學。本研究推論作者的生產力較高,同時也將有較多的合著對象。
2) 高被引的期刊與論文:在這個研究裡發現約有43%的高被引文獻是在1970到1989年間出版,顯示資訊檢索已是一個相當成熟的領域。另外,生產力較高的作者群與高被引文獻間的關連不明顯。引用最高的出版品與作者都是來自電腦科學,引用最高的前三名作者以及他們的引用次數佔全部引用數的比率分別是Salton (26.1%), Jansen (9.88%), and Baeza-Yates (8.38%)。

3)作者指定的關鍵詞(author-assigned keywords):在關鍵詞所形成的共現網路裡,除了information retrieval,中介性較高的關鍵詞還包括information seeking、information system、evaluation和user studies。另外,在網路圖上可以明顯看出這些詞語形成的四個主要的研究集群為1)使用者研究(user studies)、網路資訊檢索(Web information retrieval)、3)引用分析/科學計量學(citation analysis/scientometrics)、和4)資訊檢索系統評估(information retrieval system evaluation)。
4) 活躍的機構(active institutions):資訊檢索領域的主要研究機構大部分位於美國,並且絕大多數是學術機構或大學。利用作者的合著關係將這些研究機構的合作情形畫成網路圖,結果發現這個研究領域相當鼓勵跨機構和跨國的研究。

5) 來自於其他領域的想法(the import of ideas from other disciplines):從論文的引用關係發現,資訊檢索研究的想法主要來自以下的五個領域:電腦科學(computer science)、圖書資訊學(library and information science)、工程(engineering)、電信傳播(telecommunications)和管理(management),高達91.6%的引用來自這五個領域。
We analyzed citation data in the information retrieval subfield for the past decade (2000–2009) and presented the results in terms of co-authorship network, highly productive authors, highly cited journals and papers, author-assigned keywords, active institutions, and the import of ideas from other disciplines.
An indicator of the maturity of an area of inquiry is the growth in the number and quality of research publications (Van den Beselaar & Leydesdorff, 1996). Insight into the nature of a field can be gained by examining ‘‘the publications produced by its practitioners. To the extent that practitioners in the field publish the results of their investigations, this mode for assessing the state of a field can reflect with great specificity the content and problem orientations of the group. Of the many ways that publications can be analyzed and counted, perhaps the most revealing kind of data are the references cited by the practitioner group in their publications’’ (Small, 1981, p. 39).
A field, discipline, or other area of study can be broadly divided into an intellectual base and current research fronts (Chen, 2006). ‘‘If we define a research front as the state of the art of a specialty (i.e., a line of research), what is cited by the research front forms its intellectual base’’ (Chen, 2006, p. 360). Previous literature of a discipline cited in its current publications (i.e., its intellectual base) can inform us about current research fronts. It is through references to sources that authors make connections between concepts (Small, 1981). Collectively, such connections create a ‘‘representation of the cognitive structure of the research field’’ (Small, 1981, p. 39).
Although a number of studies have examined the literature of library and information science (e.g., Åström, 2007, 2010; Chen, Ibekwe-SanJuan, & Hou, 2010; Cronin & Meho, 2007; Donohue, 1972; Harter & Hooten, 1992; Harter, Nisonger, & Weng, 1993; Persson, 1994; Rice, 1990; White & McCain, 1998; Zhao & Strotmann, 2008a, 2008b), the nature of the literature concerning information retrieval has not been so thoroughly investigated (Ding, Chowdhury, & Foo, 2000; Ding, Yan, Frazho, & Caverlee, 2009; Ellis, Allen, & Wilson, 1999).
According to Chen (2006), the trends and patterns of scientific literatures provide researchers or communities of similar interests an overview of the related field(s) and relationships among the specific fields. More specifically, such information as the most influential articles or books, the evolvement of terms, noun phrases, keywords, the most reputed researchers, the connection between different institutions and countries over time can show trends and patterns that provide more overview.
The most prominent conclusions address the stable, multidisciplinary nature of the field. For instance, Persson (1994) found that, for a 5-year period (1986–1990), the intellectual base of the Journal of the American Society for Information Science (consisting of the most frequently co-cited authors) and its research front (consisting of articles sharing at least five cited authors) had similar maps (depicted in two-dimensional spaces). This finding highlights the stable nature of the topics explored by the information studies field.
Ding et al. (2000) conducted one of the earliest studies on the literature of information retrieval specifically. They analyzed co-citation data of 50 highly cited journals using multidimensional scaling and cluster analysis. They produced two-dimensional maps of the structure of the literature of information retrieval for an 11-year period (1987–1997). The visualizations revealed strong relationships between information retrieval and five other disciplines: computer science, physics, chemistry, psychology, and neuroscience. Their maps also show information retrieval as being part of both the computer science and LIS fields.
Tang (2004) identified the most common disciplines to which LIS exports ideas (based on the number of citations it received from the disciplines):
computer science, communication, education, management, business, and engineering. Another study looked at the export and import of ideas to and from LIS and found it to be an exporter of ideas (Cronin & Meho, 2008). That contrasts sharply with the state of the field 20 years ago when few researchers from other disciplines cited LIS literature (Cronin & Pearson, 1990).
Result
Analysis of authorship and co-authorship is critical to the understanding of scholarly communication and knowledge diffusion (Chen, 2006).
To show the extent of collaboration by the most productive researchers in our dataset, we used CiteSpace to create a coauthorship network map. Two researchers have co-authorship if they have co-authored at least one paper together. CiteSpace generates networks by measuring betweenness centrality scores. ‘‘In a network, the betweenness centrality of a node measures the extent to which the node plays a role in pulling the rest of the nodes in the network together. The higher the centrality of a node, the more strategically important the node is.’’ (Chen et al., 2009, p. 236).
Two of the author co-citation clusters are large enough to contain at least eight members/authors. Two highly productive authors (see Table 1) anchor the two clusters: Jarvelin K and Chen HC. ... These results point to a relation between collaboration frequency and the most productive authors. We are tempted to conclude that the more actively an author collaborates, the more productive she or he is. Further research is necessary to confirm this assertion.

Chen, Cribbin, Macredie, and Morar (2002) showed that visualization can be used to track the development of a scientific discipline and present the long-term process of its competing paradigms. They also assert that, among a discipline’s co-cited publications, the cluster consisting of the most highly cited publications may represent the discipline’s core or predominant paradigm.
A product of CiteSpace, Fig. 2 displays a document co-citation network, generated from the collective citing behavior in our information retrieval dataset. The network is composed of 121 reference nodes and 1163 co-citation links.

Author-assigned keywords can reveal specific focus areas of research in a field.
As suggested by Table 4, ‘‘information retrieval’’ is located near the center of the cloud of keywords. Fig. 4 also indicates that other related terms have taken on a central role in the subfield: ‘‘information seeking,’’ ‘‘information system,’’ ‘‘evaluation,’’ and ‘‘user studies.’’ This highlights the rising emphasis on user-centered system design and retrieval, as well as the importance of user studies in the evaluation of IR systems. Information retrieval research is stronger today because it has increasingly focused on user-centered design. Current user studies research is more about ‘‘users’ interaction with information retrieval systems than about user information behavior in general’’ (Zhao & Strotmann, 2008a, p. 2077).
Fig. 4 also suggests that the information retrieval subfield has its own special areas of inquiry. Four main clusters can be discerned on the visualization map, centered around user studies, Web information retrieval, citation analysis/scientometrics, and information retrieval system evaluation.
Fig. 5 maps the collaboration between the top 20 institutions of information retrieval authors in our dataset. ... These diverse groupings indicate that the information retrieval subfield encourages collaboration across institutions and countries.
Information retrieval researchers in our dataset cite primarily computer science and library and information science publications (see Table 6). Those two fields account for 82.79% of the citations. ... Apart from LIS and computer science, the third, fourth, and fifth other disciplines from which information retrieval imports ideas are engineering, telecommunications, and management, respectively. In fact, 91.6% of the citations by information retrieval authors whose articles were published between 2000 and 2009 were to these five disciplines.