顯示具有 diachronic research 標籤的文章。 顯示所有文章
顯示具有 diachronic research 標籤的文章。 顯示所有文章

2015年3月30日 星期一

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

確認科學領域的專業(specialties)本質是資訊科學的一項基本挑戰 (Morris & Van der Veer Martens, 2008; Tabah, 1999) 。由於1)可取用的書目資料來源愈來愈普及;2)網路上愈來愈多可提供分析與視覺化的電腦軟體工具;3)從多元來源而大量的資料吸收的要求愈來愈劇烈等原因,因此有愈來愈多的相關研究。共被引分析是對科學進行量化分析最常用的方法之一,特別是作者共被引分析 (author cocitation analysis, ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b)以及文件共被引分析 (document cocitation analysis, DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985)。作者共被引分析的目的在透過被相關文獻一起引用的作者群集,確認領域裡的專業。重要的作者共被引分析研究包括White & McCain (1998),這個研究以1972到1995年間12種資訊科學相關期刊的120位高被引作者進行作者共被引分析,研究結果發現當時的資訊科學分為兩個基本上彼此獨立的陣營:資訊檢索(information retrieval)與文獻(literature)。Zhao and Strotmann (2008a, 2008b) 以1996-2005年的資訊科學相關期刊資料重新進行了相同的研究,他們的結果發現了5個主要的專業:使用者研究(user studies)、引用分析(citation analysis)、實驗型檢索(experimental retrieval)、網路計量學 (Webometrics)以及知識領域的視覺化(visualization of knowledge domains),其中新興的兩個專業:網路計量學和知識領域的視覺化連繫了引用分析以及實驗型檢索,而使用者研究則是此時最大的專業。Aström (2007) 則是使用文件共被引分析的例子,他們分析了1990到2004年的21種圖書資訊學期刊,利用多維尺度法(multidimensional scaling, MDS)產生結果,他們的結果與White & McCain (1998)的研究類似,整個領域可分為兩個陣營,不過Aström (2007)的結果將稱為資訊尋求與檢索(information seeking and retrieval),而不是資訊檢索。

不管是作者共被引分析或是文件共被引分析其步驟大致如下:
1) 檢索引用資料。
2) 建構參考文件或作者共同被引用的矩陣。
3) 將共被引矩陣表示成節點與連結的圖(node-and-link graph)或是多維尺度法的組態(configuration),並且可以利用尋路網路(Pathfinder network scaling)或最小生成樹(minimum spanning tree)裁減連結。
4) 利用群集、社群發現(community finding)、因素分析(factor analysis)、主成分分析(principle component analysis)或者隱含語意索引(latent semantic indexing)等各種演算法確認專業。例如Morris & Van der Veer Martens (2008)、 Persson (1994)、 Tabah (1999)、 White & Griffith (1982)以及Janssens, Leta, Glänzel, and De Moor (2006)。
5) 根據群集成員間共同的主題(themes),解釋共被引群集的性質。通常需要豐富的領域知識,而且是一個花費大量時間與認知需求(cognitively demanding)的工作。

本研究對於作者共被引以及文件共被引形成的群集進行結構與動態的描述與解釋,分析的資料為1996到2008年間的12種資訊科學(information science)領域相關期刊,共計10853筆書目紀錄,引用的參考文獻為129060筆,引用次數為206180,而參考文獻的作者共有58711位。本研究以餘弦(cosine)測量作者或文件之間的關連大小,做為節點間的連結,建立網路;然後計算從原先網路導出的Laplacian矩陣(Laplacian matrices)的特徵向量(eigenvectors)找出群集。這種利用標準線性代數的頻譜群集(spectral cluster)演算法,較其他的群集演算法更有效率,而且因為不需要假設群集的形式,所以更有彈性與強健。標註群集方面則是利用引用文獻論文的詞語與摘要句,詞語包括題名與摘要中出現的名詞片語與索引詞(index terms),利用 tf*idf (Salton, Yang, & Wong, 1975)、對數似然比(log-likelihood ratio, LLR)測試 (Dunning, 1993)以及相互資訊(mutual information, MI)等三種資訊做為判斷的參考。摘要句則是從題名與摘要尋找最有代表性的句子,例如以Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007)對句子進行排序。




A multiple-perspective cocitation analysis method is introduced for characterizing and interpreting the structure and dynamics of cocitation clusters.

The generic method is applied to a three-part analysis of the field of information science as defined by 12 journals published between 1996 and 2008: (a) a comparative author cocitation analysis (ACA), (b) a progressive ACA of a time series of cocitation networks, and (c) a progressive document cocitation analysis (DCA).

Identifying the nature of specialties in a scientific field is a fundamental challenge for information science (Morris & Van der Veer Martens, 2008; Tabah, 1999).

The growing interest in mapping and visualizing the structure and dynamics of specialties is because of a number of reasons:
1. Widely accessible bibliographic data sources such as the Web of Science, Scopus, and Google Scholar (Bar-Ilan, 2008; Meho & Yang,2007) as well as domain-specific repositories such as ADS (http://www.adsabs.harvard.edu/) and arXiv (http://arxiv.org/).
2. Freely available computer programs and Web-based general-purpose visualization and analysis tools such as ManyEyes (http://manyeyes.alphaworks.ibm.com/) and Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/; Batagelj & Mrvar, 1998), special-purpose citation analysis tools such as CiteSpace (http://cluster.cis.drexel.edu/&u0007E;cchen/citespace/; Chen, 2004; Chen, 2006), and social network analysis such as UCINET (http://www.analytictech.com/ucinet6/ucinet.htm).
3. Intensified challenges for digesting the vast volume of data from multiple sources (e.g., e-Science, Digging into Data (http://www.diggingintodata.org/), cyber-enabled discovery, SciSIP; Lane, 2009).

Cocitation studies are among the most commonly used methods in quantitative studies of science, especially including author cocitation analysis (ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b) and document cocitation analysis (DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985).

For instance, once cocitation clusters are identified, assigning the most meaningful labels for these clusters is currently a challenging task because any representative labels of clusters must characterize not only what clusters appear to represent, but also the salient and unique reasons for their formation.

The new procedure reduces analysts' cognitive burden by automatically characterizing the nature of a cocitation cluster in terms of (a) salient noun phrases extracted from titles, abstracts, and index terms of citing articles and (b) representative sentences as summarizations of clusters.

ACA aims to identify underlying specialties in a field in terms of groups of authors who were cited together in relevant literature.

White & McCain (1998) presented a comprehensive view of information science based on 12 journals in library and information science across a 24-year span (1972–1995). It analyzed cocitation patterns of 120 most-cited authors with factor analysis and multidimensional scaling. The authors drew upon their extensive knowledge of the field and offered an insightful interpretation of 12 specialties identified in terms of 12 factors. The most well-known finding of the study is that information science at the time consisted of two essentially independent camps, namely, the information retrieval camp and the literature camp, including citation analysis, bibliometrics, and scientometrics.

Zhao and Strotmann (2008a, 2008b) followed up White and McCain's study using the same set of 12 journals and the same number of 120 cited authors in an updated time frame of 1996-2005. ... Zhao and Strotmann (2008b) found five major specialties and manually labeled them as user studies, citation analysis, experimental retrieval, Webometrics, and visualization of knowledge domains. In contrast to the findings of (White & McCain, 1998), experimental retrieval and citation analysis retained their fundamental roles in the field, and the user studies specialty became the largest specialty. Webometrics and visualization of knowledge domains appeared to make connections between the retrieval camp and the citation analysis camp.

A DCA by Aström (2007) studied papers published between 1990 and 2004 in 21 library and information science journals. Results were depicted in multidimensional scaling (MDS) maps. Aström's study also identified the two-camp structure found by (White & McCain, 1998). On the other hand, Aström found an information seeking and retrieval camp, instead of the information retrieval camp as in (White and McCain).

Although manually labeling a cocitation cluster can be a very rewarding process of learning about the underlying specialty and result in insightful and easy to understand labels, it requires a substantial level of domain knowledge and it tends to be time-consuming and cognitively demanding because of the synthetic work required over a diverse range of individual publications.

Traditionally, researchers often identify the nature of a cocitation cluster based on common themes among its members. ... The emphasis on common areas is a practical strategy; otherwise, comprehensively identifying the nature of a specialty can be too complex to handle manually.

Many researchers have studied the structural and dynamic properties of specialties in information science in terms of clusters, multivariate factors, and principle components (Morris & Van der Veer Martens, 2008; Persson, 1994; Tabah, 1999; White & Griffith, 1982).

A recent study of information science (Ibekwe-SanJuan, 2009) mapped the structure of information science at the term level using a text analysis system TermWatch and a network visualization system Pajek, but it did not address structural patterns of cited references.

Researchers also studied the structure of information science qualitatively, especially with direct inputs from domain experts. For example, Zins conducted a Critical Delphi study of information science, involving 57 leading information scientists from 16 countries (Zins, 2007a, 2007b, 2007c, 2007d).

Janssens, Leta, Glänzel, and De Moor (2006) studied the full-text of 938 publications in five library and information science journals with latent semantic analysis (LSA; Deerwester, Dumais, Landauer, Furnas, & Harshman, 1990) and agglomerative clustering. They found an optimal 6-cluster solution in terms of a local maximum of the mean silhouette coefficients (Rousseeuw, 1987) and a stability diagram (Ben-Hur, Elisseeff, & Guyon, 2002). Their clusters were labeled with single-word terms selected by tf*idf (p. 1625), which are not as informative as multiword terms for cluster labels.

Klavans, Persson, and Boyack (2009) recently raised the question of the true number of specialties in information science. They suspected that the number is much more than the 11 or 12 as reported in ACA studies such as (White & McCain, 1998) and (Zhao & Strotmann, 2008a, 2008b), but significantly fewer than the 72 reported in their own study, which is also based on the 12 journals between 2001 and 2005.

The 12-journal Information Science dataset, retrieved from the Web of Science, contains 10,853 unique bibliographic records, written by 8,408 unique authors from 6,553 institutions and 89 countries. These articles cited 129,060 unique references for a total of 206,180 times. They cited 58,711 unique authors and 58,796 unique sources.

The traditional procedure of cocitation analysis for both DCA and ACA comprises the following steps:
1. Retrieve citation data from sources such as the Science Citation Index (SCI), Social Science Citation Index (SSCI), Scopus, and Google Scholar.
2. Construct a matrix of cocited references (DCA) or authors (ACA).
3. Represent the cocitation matrix as a node-and-link graph or as a multidimensional scaling (MDS) configuration with possible link pruning using Pathfinder network scaling or minimum spanning tree algorithms.
4. Identify specialties in terms of cocitation clusters, multivariate factors, principle components, or dimensions of a latent semantic space using a variety of algorithms for clustering, community finding, factor analysis, principle component analysis, or latent semantic indexing.
5. Interpret the nature of cocitation clusters.

The interpretation step is the weakest link. It is time-consuming and cognitively demanding, requiring a substantial level of domain knowledge and synthesizing skills. In addition, much of attention routinely focuses on cocitation clusters per se, but the role of citing articles that are responsible for the formation of such cocitation clusters may not be always investigated as an integral part of a specialty.

Our new method extends and enhances traditional cocitation methods in two ways: (a) by integrating structural and content analysis components sequentially into the new procedure and (b) by facilitating analytic tasks and interpretation with automatic cluster labeling and summarization functions. The new procedure is highlighted in yellow in Figure 2, including clustering, automatic labeling, summarization, and latent semantic models of the citing space (Deerwester et al., 1990).

Our new procedure adopts several structural and temporal metrics of cocitation networks and subsequently generated clusters.

Structural metrics include betweenness centrality, modularity, and silhouette.

Temporal and hybrid metrics include citation burstness and novelty

The betweenness centrality metric is defined for each node in a network. It measure the extent to which the node is in the middle of a path that connects other nodes in the network (Brandes, 2001; Freeman, 1977). High betweenness centrality values identify potentially revolutionary scientific publications (Chen, 2005) as well as gatekeepers in social networks.

In the context of this study, the modularity Q measures the extent to which a network can be divided into independent blocks, i.e., modules (Newman, 2006; Shibata, Kajikawa, Taked, & Matsushima, 2008).

The silhouette metric (Rousseeuw, 1987) is useful in estimating the uncertainty involved in identifying the nature of a cluster.

Burst detection determines whether a given frequency function has statistically significant fluctuations during a short time interval within the overall time period.

Sigma is introduced in (Chen, et al., 2009a) as a measure of scientific novelty. ... In this study, Sigma is defined as (centrality + 1)burstness such that the brokerage mechanism plays more prominent role than the rate of recognition by peers.

We adopt a hard clustering approach such that a cocitation network is partitioned to a number of nonoverlapping clusters.

In this article, cocitation similarities between items i and j are measured in terms of cosine coefficients.

A good partition of a network would group strongly connected nodes together and assign loosely connected ones to different clusters. This idea can be formulated as an optimization problem in terms of a cut function defined over a partition of a network. Technical details are given in relevant literature (Luxburg, 2006; Ng, Jordan, & Weiss, 2002; Shi & Malik, 2000).

Spectral clustering is an efficient and generic clustering method (Luxburg, 2006; Ng et al., 2002; Shi & Malik, 2000). It has roots in spectral graph theory. Spectral clustering algorithms identify clusters based on eigenvectors of Laplacian matrices derived from the original network.

Spectral clustering has several desirable features compared to traditional algorithms such as k-means and single linkage (Luxburg, 2006):
 • It is more flexible and robust because it does not make any assumptions on the forms of the clusters,
• it makes use of standard linear algebra methods to solve clustering problems, and
• it is often more efficient than traditional clustering algorithms.

Candidates of cluster labels are selected from noun phrases and index terms of citing articles of each cluster. These term are ranked by three different algorithms. In particular, noun phrases are extracted from titles and abstracts of citing articles. The three term ranking algorithms are tf*idf (Salton, Yang, & Wong, 1975), log-likelihood ratio (LLR) tests (Dunning, 1993), and mutual information (MI).

Each cocitation cluster is summarized by a list of sentences selected from the abstracts of articles that cite at least one member of the cluster.

In this study, sentences are ranked by Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007). Given a set S of N sentences, let M be the square matrix that for each pair of sentences gives the number of nominal words in common (nouns and adjectives).

In this study, summarization sentences were also ranked by two new functions gtf and gftidf , which are further simplified approximations of the energy function E.

The ACA and DCA studies described in this article were conducted using the CiteSpace system (Chen, 2004; Chen, 2006). CiteSpace is a freely available Java application for visualizing and analyzing emerging trends and changes in scientific literature.

CiteSpace supports a unique type of cocitation network analysis—progressive network analysis—based on a time slicing strategy and then synthesizing a series of individual network snapshots defined on consecutive time slices. Progressive network analysis particularly focuses on nodes that play critical roles in the evolution of a network over time. Such critical nodes are candidates of intellectual turning points.

In summary, (a) spectral clustering and factor analysis identified about the same number of specialties, but they appeared to reveal different aspects of cocitation structures and (b) cluster labels chosen from citers of a cluster tend to be more specific terms than those chosen by human experts.

We found the comparison with the study of Zhao and Strotmann very valuable. It offered us an opportunity to compare the analysis conducted by human experts to the interpretation cues provided by our automatic labeling and summarization methods.

Spectral clustering for the purpose of network decomposition is exclusive in nature although in reality it is often sensible to allow overlapping clusters because of multiple roles individual entities may play.

Spectral clustering of cocitation networks tends to generate distinct clusters with high precision, whereas human experts tend to aggregate entities into broadly defined clusters.

In conclusion, the new cocitation analysis procedure has the following advantages over the traditional one:
• It can be consistently used for both DCA and ACA.
• It uses more flexible and efficient spectral clustering to identify cocitation clusters.
• It characterizes clusters with candidate labels selected by multiple ranking algorithms from the citers of these clusters and reveals the nature of a cluster in terms of how it has been cited.
• It provides metrics such as modularity and silhouette as quality indicators of clustering to aid interpretation tasks.
• It provides integrated and interactive visualizations for exploratory analysis.

Modularity and silhouette metrics provide useful quality indicators of clustering and network decomposition.

2014年9月15日 星期一

Bonnevie-Nebelong, E. (2006). Methods for journal evaluation: journal citation identity, journal citation image and internationalization. Scientometrics, 66(2), 411-424.

Bonnevie-Nebelong, E. (2006). Methods for journal evaluation: journal citation identity, journal citation image and internationalization. Scientometrics, 66(2), 411-424.

Scientometrics

本研究以引用分析對Journal of Documentation (J DOC) 進行評估,並且與JASIST和JIS進行比較。所使用的引用分析方法包括三個方面:以引用的參考文獻為主的期刊引用認同(journal citation identity)、以被引用的情形為主的期刊引證形象(journal citation image)和以出版品本身為主的國際化(internationalisation)。

在期刊引用認同方面有兩種指標。第一種指標是引用對被引用者比(citations/citee-ratio),計算方式是分析範圍內所有參考文獻數除以參考文獻上出現的期刊種類,如果這個數值愈低,表示出現許多不同種類的期刊,也就是使用的期刊具有多元性(diversity)。另一個指標是自我引用(self-citations),用來測量期刊在科學領域內的獨立性(isolation),如果自我引用的程度低表示在科學領域內的影響力高。自我引用指標的測量包括引用文獻中來自本身期刊的比例(self-citing)和期刊被引用的情形下來自本身的比例(self-cited),前者是期刊引用認同的一部份,而後者則屬於期刊引證形象。

期刊引證形象也包含兩種指標。第一種指標是新期刊擴散因素(new journal diffusion factor),此一指標分析該期刊時間範圍內每一篇論文平均被引用的期刊種類,代表該期刊的想法出口情形(export of ideas)、跨領域性(transdisciplinarity)以及專殊化(specialisation)程度。另一只標示該期刊的共被引期刊,根據共被引情形以及共被引期刊的期刊影響因素(Journal Impact factor)來加以描述。

國際化是測量出版品以及引用期刊論文的作者地區。

各種分析方法與指標整理為Table 1。



首先,引用對被引用者比的結果如Figure 1。另外,1990到2003年的平均引用對被引用者比,JDOC為1.50,JASIST為1.88,JIS則為1.44。較低的引用對被引用者比表示引用的文獻裡重複的期刊較多種,代表這份期刊有較多元的科學基礎(scientific base)。從結果上看來,JDOC比JASIST的科學基礎多元性較高,但較JIS來得低。

JDOC比JASIST和JIS的文章有較高的比例是書評(boo review),這使得JDOC的參考文獻數較少,因為書評平均只有1.6到2筆參考文獻。

在1990到2003年間,JDOC、JASIST和JIS等三種期刊引用本身的比例都有下降的趨勢,表示這三種期刊愈來愈不孤立,測量期刊被引用的情形,則可發現JDOC與JIS來自期刊本身的比例則較低,表示它們在這個領域的能見度(visibility)較高。另外,JIS在從1979年開始的前十年引用來自期刊本身的比例較高,則說明了這個期刊在當時為在領域邊緣的新期刊。

JDOC的新期刊擴散因素比其他兩種期刊稍大,並且有往上的趨勢。

經常與JDOC共同引用的前十種期刊如Table 3所示。期刊共被引的相似度以Jaccard Similarity測量。

JDOC上論文作者的地區分布如Figure 10。主要的作者來自西歐地區,並逐漸增加。



引用JDOC論文的作者地區分布則如Figure 11。以北美地區的作者引用最多,但西歐地區則逐漸增加。



The Journal Citation Identity is a reference analysis. It is measured by looking into  the referencing style of the publishing authors. What is their combined citations/(journal) citee-ratio? This means that the total number of references in the journal must be calculated, year-by-year or all years taken together. The result of this is divided with the number of different journals present in the set of references. If the set contains many different journals, the ratio will be lower. Consequently a low average signifies a greater diversity in the use of journals among the authors as part of their scientific base, and thus a wider horizon.

Self-citations are part of the Journal Citation Identity as well as the Journal Citation Image, depending on the perspective. ... They are indicators of the style of a journal. Many self-citations among the references may signify isolation of the journal in the scientific domain (high rate). A low rate of self-citations may indicate a high level of influence in the scientific community.

The Journal Citation Image is based on citation analyses of two types: the New Journal Diffusion Factor (N JDF) and journal co-citation analysis.

The New Journal Diffusion Factor was proposed by Frandsen, and is inspired by Rowlands’ diffusion factor. It measures breadth by number of citing journals per published article. N JDF is the average number of different journals that an average article is cited by within a given time window. The result of this tells about the scientific style and about breadth, export of ideas, transdisciplinarity and degree of specialisation of a journal. N JDF is tested for JDOC in a time perspective.

The Journal Citation Image “the White way” means to do a co-citation journal-by-journal analysis and interpret the result in a qualitative manner. It is thus a means to evaluate a journal by the journals co-cited with the journal in question. The co-cited journals are displayed in a list ranked by frequency of co-incidences, the number of citations for each co-cited journal taken into consideration by application of the jaccard calculations. Also the Journal Impact factor (JIF) is used to evaluate the co-cited journals. The co-cited journals then function as image-makers of the journal in question.

Internationalisation is measured by looking into the geographic locations of both publishing and cited authors of the JDOC.

A high citation/citee ratio means that the journal has many recited journals among its references. A low ratio signifies less journal re-citations and thus a greater diversity of journals as part of the scientific base and a wider horizon among authors.

Journal self-citations. Journal self-citations can be analysed from two perspectives, by self-citing rate and by self-cited rate. The first mentioned is part of the citation identity, the second one is part of the self-image, but the two types of self-citations are treated together here for practical reason.

The three journals all show decreasing self-citing rates during the years 1980–2003. This may signify a tendency towards less isolation of the field.

2014年9月9日 星期二

Åström, F. (2007). Changes in the LIS research front: Time‐sliced cocitation analyses of LIS journal articles, 1990–2004. Journal of the American Society for Information Science and Technology, 58(7), 947-957.

Åström, F. (2007). Changes in the LIS research front: Time‐sliced cocitation analyses of LIS journal articles, 1990–2004. Journal of the American Society for Information Science and Technology, 58(7), 947-957.

scientometrics

本研究利用論文間的共被引分析探討1990到2004年間圖書資訊學(LIS)的研究前沿(research front)的改變,了解這個學科目前的處境與發展趨勢。分析資料為21種LIS期刊。將同時間內具有影響力的共被引文章定義為研究前沿(research fronts),並且分為三個5年期間,分析領域的改變。研究結果發現LIS由兩個不同研究領域構成的穩定結構:資訊計量學(informetrics)和資訊搜尋與檢索(information seeking and retrieval),由於分享研究興趣與方法,資訊檢索與資訊計量學有靠近的傾向。而網路為主的研究成為資訊計量學和資訊搜尋與檢索的主要研究則是這個領域的主要變化。

本研究採用的期刊來源為JCR (2003) 的Information Science & Library Science分類下的55種期刊。去除主要是被非LIS期刊引用的期刊以及評論性或商業性期刊後,選擇1990到2004年間有出版的期刊,如下表共21種。

論文的總數為13605筆,從中選取最高被引用的論文,建立共被引次數矩陣。以多維尺度演算法(multidimensional scaling algorithm, MDS)進行處理。
首先是研究基礎(research base)部分,從13605筆論文資料的221586次引用(150145篇參考文獻)中,選取被引用超過50次的文獻,共66筆進行分析。其共被引映射圖如FIG 1.:

與先前研究一致,圖書資訊學在圖形上分為兩個區域,圖形上半部為資訊搜尋與檢索相關論文,下半部則為資訊計量學文獻,此一結果和 Persson (1994) 與 White & McCain (1998)等研究相符合。另外在資訊計量學文獻右邊,還有一群文獻形成網路計量學(webometrics)叢集。網路計量學是利用連結、引用與叢集等資訊計量學方法進行網路本質與特性的分析。

資訊搜尋與檢索從早期的系統導向資訊檢索(systems-oriented information retrieval)發展到使用者-系統互動研究(user-system interaction studies)和資訊行為(information behavior)。

除了網路計量學以外,資訊計量學以書目計量映射(bibliometric mapping)為中心,周圍的部分是書目計量分布(bibliometric distributions)。

為了進一步了解與核對共被引分析的結果,將共被引資料輸入叢集分析。叢集分析所產生的8個叢集符合映射圖的結構,各叢集如TABLE 2。
從TABLE 2各叢集出版年度的中位數,可以將八個叢集分為四個時期:第一個時期圖書資訊學的研究包括實驗性資訊檢索(experimental information retrieval)、書目計量映射以及書目計量分布;第二個時期開始對於資訊檢索的使用者端產生興趣,增加了搜尋過程與認知面向的資訊檢索研究;隨後是在1990年代早期進行的相關性(relevance)研究,同時也傾向於一般的資訊行為;1990年代末期則受到網路科技的影響,開始進行網路以及網路計量學的研究。

接下來的共被引分析,被引用的參考文獻僅限於也在13605篇論文裡的論文,來了解具有影響力的論文,做為研究脈絡(research context)。選取被引用次數超過25次的論文,共65篇。呈現的圖形大致上仍然可明顯的看出分為上半區域的資訊搜尋與檢索和下半區域的資訊計量學。但資訊檢索的研究以認知性資訊搜尋與檢索、相關性和資訊行為為主要,實驗性資訊檢索研究則成為邊緣。

相較於研究基礎,在研究脈絡上可以發現資訊計量學的結果較為分散,包含三個部分:研究合作(research collaboration)、書目計量映射與網路計量學,並且以網路計量學最為主要。


以TABLE 3的叢集結果來看,在研究脈絡中雖然實驗性資訊檢索與書目計量分布消失了,但增加了兒童的資訊行為研究和對於研究合作的資訊計量學分析。雖然這些研究依然存在,但本身並沒有形成叢集,而是歸入其他的叢集中,如IR/Search。

對三個5年的時期進行研究前沿分析,第一個時期1990-1994年,共有3401篇論文,彼此間有1581次引用,39篇論文獲得5次以上的引用。這個時期以ISR為主,特別是使用者觀點的ISR研究;資訊計量學由兩個小叢集組成:一為研究合作,另一聚焦於映射。

第二個時期1995-1999年,包含3318篇論文,彼此間的引用共有2117次,獲得5次以上引用的論文共有52篇。這時期ISR的聚集相當明顯,除了聚焦在資訊科技(information technology)和實驗性資訊檢索(experimental IR)的兩個叢聚外。在一般的資訊計量學之外,另外還有研究成效(research performance)的叢集。

第三個時期2000-2004年,有4147筆論文,彼此間有2926次的引用,62篇論文的引用次數超過7次。在這個時期,可以看出資訊計量學較前面兩個時期緊密連接,主要聚焦在網路計量學,而ISR則較前兩個時期變得較為分散,可分為三個叢集:ISR、兒童的資訊行為(children's information behaviors)以及健康資訊學(health informatics)。

本研究發現LIS有相當穩定的結構,主要為ISR及資訊計量學所構成。另外,從研究基礎上發現,大多為理論或方法學的文獻,但研究脈絡與前沿上的文獻卻以實務性的論文為主。就三個時期的研究來看,1900-1994年以圖書館與資訊服務(library and information service)為主,第二個時期則是線上資料庫與資訊尋求;第三個時期受到WWW影響,主要的研究從群體利用WWW搜尋資訊的方法到發展分析網站影響因素的方法。最後,本研究發現ISR與資訊計量學有愈來愈接近的趨勢,其原因是因為兩者都需要測量文件(或搜尋問題)之間的關係強度,並且也都對將資訊視覺化有興趣,因此彼此引用整合的機會增加。

Based on articles published in 1990–2004 in 21 library and information science (LIS) journals, a set of cocitation analyses was performed to study changes in research fronts over the last 15 years, where LIS is at now, and to discuss where it is heading.

The results show a stable structure of two distinct research fields: informetrics and information seeking and retrieval (ISR). However, experimental retrieval research and user oriented research have merged into one ISR field; and IR and informetrics also show signs of coming closer together, sharing research interests and methodologies, making informetrics research more visible in mainstream LIS research. Furthermore, the focus on the Internet, both in ISR research and in informetrics—where webometrics quickly has become a dominating research area—is an important change.

The nature and intellectual organization of LIS has been thoroughly investigated in analyses describing the general traits of LIS research, as well as mapping how LIS has been organized in different research themes (Persson, 1994; White & Griffith, 1981; White & McCain, 1998).

My approach centers on the following questions. What research topics have dominated LIS during the period 1990–2004? What changes can be observed in the topics addressed over the last 15 years? Can these changes can be used to tell us something about where LIS is heading?

Most definitions of “research fronts” explain them as groups of citing articles being clustered through bibliographic coupling (e.g., Persson, 1994), and their relations to the cited documents clustered by cocitation analysis (Garfield, 1994; Morris et al., 2003; Price, 1965). Although Persson sees the current (citing) articles as the research front and the cited documents as the research base, Garfield, for example, also includes the clusters of cocited core articles into the research front.

In addition, by analyzing the co-occurrence of highly cited documents, we also get an indication on the impact of the articles, thus expanding the definition of research fronts as including influential, as well as current research.

To identify LIS research, and to select journals for the analyses, the Journal Citation Reports: JCR Social Sciences (Thomson ISI, 2003) was used. To defining LIS research, JCR’s Information Science & Library Science classification, covering 55 journals, was used.

To limit the definition, all general LIS journals were identified and the specialized ones were excluded. This was done using the “Citing Journal” field in JCR: If the journal primarily was cited by non-LIS publications, it was excluded from the study.

The analyses were done on a document level, as opposed to an analysis on the author level. Although an author analysis provides more of an overview, the document analysis is more detailed, e.g., by not grouping documents on different topics by the same author.

The result reflects contemporary and influential research within a specific field of research, i.e., the research front.

The research base was based on the 13,605 journal articles published from 1990–2004 and their 221,586 references to 150,145 unique documents. The 66 most-cited documents that received 50 citations or more were selected for further analysis (Figure 1).

The map shows two main areas consistent with the structures found in earlier analyses on LIS (e.g., Persson, 1994; White & McCain, 1998). On the top half of the map, a group of information-seeking and retrieval (ISR) related literature is featured and on the bottom half, a group of informetrics literature. However, on the right side of the informetrics field, a group of webometric studies has formed a cluster. Webometrics is the study of the nature and properties of the World Wide Web, using informetric methodologies such as link, citation, and cluster analyses (Björneborn & Ingwersen, 2001).

In the ISR section of the map, there is a thematic shift from right to left. Systems-oriented information retrieval (IR) literature is on the far right, followed towards the left by user-system interaction studies and information behavior. In comparison to Persson (1994), the “soft” part of the IR-field has increased its impact compared to the “hard” systems-oriented IR research.

Apart from the webometric group on the far right, the informetrics field is centered on bibliometric mapping, surrounded by documents concerning bibliometric distributions.

To enhance the results of the cocitation analysis, a cluster analysis (Persson, 1994) was performed, resulting in eight clusters (Table 2). The clusters support the structures identified in the map, and reveal a division of the soft IR-research: from search- and relevance-focused documents, over cognitive IR and information seeking, to information behavior.

The publication years of the clustered documents shows four generations of research orientations, a trait also visible in the IR part of the map. The first generation of LIS research includes experimental IR, bibliometric distributions, and bibliometric mapping. The second generation of research, with references published from the early 1980s marks the increasing interest in the user side of IR, incorporating the search process and the cognitive perspective into IR and LIS research. This is followed by the relevance studies in the early 1990s; and a contemporary trend to focus on general information behavior. The most recent trend in the LIS research base is studies on World Wide Web and webometrics, dating back to the late 1990s.

The results of the second analysis show influential research areas during the period 1990–2004. It is still the same 13,605 articles providing the material, but only the 18,615 citations to articles present in the set of citing documents are analyzed. Here, as well as in the following time-sliced analysis, the self-citations were removed. Out of the 5024 unique-cited documents, the 65 articles being cited 25 times or more were selected and analyzed (Figure 2).

The general structure of the map is the same: with informetrics on the lower half and ISR on the top half. There are some differences, however. In the top half, a center has developed around “Kuhlthau, 1991” and “Ingwersen, 1996,” focusing on cognitive ISR, relevance, and information behavior, while experimental IR research has become peripheral. Different perspectives on the user-oriented research has dominated the information-seeking and retrieval field; and has together with the wider information behavior field formed a strong research area of different variations on information-seeking research.

At the same time, the informetrics field has become more dispersed, with three clearly defined subfields: research collaboration to the left, bibliometric mapping in the middle, and webometrics on the right side. In comparison with the research base, webometrics has become the dominating research area within the informetrics field.

2014年8月15日 星期五

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words. Journal of the American Society for Information Science and Technology, 62(10), 1933-1953.

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words.Journal of the American Society for Information Science and Technology,62(10), 1933-1953.

Scientometrics

圖書資訊學(LIS)為對於記錄下來的資訊(recorded information)和具有文化意義的文物與標本(culturally meaningful artifacts and specimens)有興趣的研究領域(Bates, 2010),包括的領域有檔案學(archival science)、 書目(bibliography)、文獻與文類理論(document and genre theory)、資訊學(informatics)、資訊系統(information systems)、知識管理(knowledge management)、圖書資訊學(LIS)、博物館研究(museum studies)、記錄管理(records management)和資訊的社會研究(social studies of information)。過去有許多研究嘗試定義與描述圖書資訊學的領域並且確認其中包含的研究主題,這些研究使用的方法相當廣泛,包含Järvelin & Vakkari (1990, 1993)採用內容分析(content analysis);Åström (2007, 2010)、Moya-Anegón, Herrero-Solana, & Jiménez-Contreras (2006)和 Persson (1994) 針對期刊或期刊文章進行書目計量分析 (bibliometric analysis) ; Moya-Anegón et al., (2006)和White & McCain (1998)針對作者進行書目計量分析 ;Åström (2002)、 Ding, Chowdhury, & Foo (2001) 和 Janssens, Leta, Glänzel, & De Moor (2006)利用從題名、摘要或全文抽取的詞語進行詞語的共現分析(co-word analysis) ;Sugimoto & McCain (2010)則是用索引詞語的三元共現分析(tri-occurrence analysis) ; van den Besselaar & Heimeriks (2006)利用詞語和參考文獻的組合進行分析;以及Sugimoto, Li, Russell, Finlay, & Ding, (2011)和 Sugimoto & McCain (2010)所使用的主題模型分析方法。

上述的這些方法,許多必須依賴於作者對於領域知識的了解,才能了解領域的主題與認知結構(cognitive structure),例如White & McCain (1998)基於最重要的作家的集群,觀察資訊科學由圍繞在一個微弱中心的許多專業所組成;Åström (2010)則是透過作者與期刊的映射圖說明這個領域的圖書館學(LS)和資訊科學(IS)之間具有差距。除了是認知結構較不直接的指標之外,引用分析另一個的問題是不同的次領域有不同的發表與引用實務。

論文題名包含許多能夠指出該文章內容的詞語(Buxton & Meadows, 1977; Meadows, 1998)。因此,本研究採用的方法是利用期刊論文題名上的重要詞語進行分析。分析的資料來自16種LIS期刊於1988到2007年發表的10344筆論文資料。

選取100個最常出現於題名的詞語。

本研究使用的分析技術包含詞語的相對頻率(relative frequency)並且根據詞語的共現進行叢集,最後並將詞語以及期刊與發表年度等進行多維尺度分析(multidimensional scaling, MDS),產生視覺化的結果。

詞語的共現分析以及階層式集群分析的結果發現三個主要分類LS(圖書館學)、IS(資訊科學)、SCI-BIB(科學計量學-書目計量學)以及兩個較小的分類資訊尋求行為(information-seeking behavior)和書目指導(bibliographic instruction)。LS可再細分為學術圖書館專業(academic librarianship)、公共圖書館專業(public librarianship) (包含館藏建立)、資訊素養和學校圖書館專業(information literacy and school librarianship, technology)、政策(policy)、全球資訊網(the web)、知識管理(knowledge management)、數位圖書館(digital libraries)、電子商務(e-commerce)、法律(law)以及學術出版(scholarly publishing)等主題。IS則包含資訊檢索(information retrieval)、網路搜尋(web search)、分類目錄(catalogs)以及資料庫(database)等主題。SCI-BIB也有書目計量指標(bibliometric indicators)、作者生產力(author productivity)與引用研究(citation study)等主題。整體的結構如下圖

從詞語的使用可以發現LIS中有某些持續出現的核心詞語,但也有一些詞語的使用在20年間有明顯的變化,這些都是與科技相關的(technologically related)詞語,這個現象符合Saracevic(1999)所宣稱的LIS是個科技驅動的(technology driven)領域。大致上來說,LIS內的改變可以從資料庫(database),到數位圖書館(digital libraries),到全球資訊網(the World Wide Web)等詞語使用的移轉上看得出來。

除了科技驅動的特徵外,LIS同時也有很大的範圍在討論資訊尋求行為,這是LS和IS都共同關心的課題。

A number of empirical studies of LIS have been conducted with the aim of describing and defining the field and identifying research areas within it. These studies applied a wide array of approaches: content analysis (Järvelin & Vakkari, 1990, 1993); bibliometric analysis of journals and journal articles (Åström, 2007, 2010; Moya-Anegón, Herrero-Solana, & Jiménez-Contreras, 2006; Persson, 1994); bibliometric analysis of authors (Moya-Anegón et al., 2006,White & McCain, 1998); co-word analysis of both index terms and words extracted from titles, abstracts, and full text (Åström, 2002; Ding, Chowdhury, & Foo, 2001; Janssens, Leta, Glänzel, & De Moor, 2006); tri-occurrence analysis of index terms (Sugimoto & McCain, 2010); analysis of word-reference combinations (van den Besselaar & Heimeriks, 2006); and topic analysis (Sugimoto, Li, Russell, Finlay, & Ding, 2011; Sugimoto & McCain, 2010).

Some notable studies of cognitive structure of LIS have interpreted topics post hoc, by assigning topicality based on knowledge of the author’s domain (e.g., White & McCain, 1998). In White and McCain’s influential visualization of LIS, they concluded that “information science lacks a strong central author, or group of authors, whose work orients the work of others across the board. The field consists of several specialties around a weak center” (p. 343). However, this analysis was based foremost on the clustering of authors, rather than topics. Similarly, Åström (2010) examined the divide between LS and IS components of the field by a bibliometric mapping of authors and journals. Topicality was assigned through expert knowledge of the domains in which these authors wrote and journals published.

Of the various components of textual documents, the titles, and the choice of words in them, are of particular importance. Title words function as “attention triggers” (Bazerman, 1985, 1988). They are devices for capturing interest in the world where information overload is a norm. Title words
have been called “signal-words”1 (Rip & Courtial, 1984) and “macro-actors” or “macro-terms”2 (Callon et al., 1983). Titles of journal articles themselves have undergone a change during the 20th century, becoming more informative, more specific, and containing a larger number of words that indicate article content (Buxton & Meadows, 1977; Meadows, 1998). Leydesdorff (1989) claims that “title words seem to offer a means of making visible the internal cognitive structure” (p. 217) of a discipline. He also claims that “word structure reflects internal intellectual organization in terms
of the codification of word usage in the relevant disciplines” (Leydesdorff, 1989, p. 221). 

Co-word analysis is based on co-occurrence of words (all words, or selected keywords) extracted from titles, abstracts, or text in general, or the index terms assigned by authors or indexers. Co-word analysis is a method that derives “higher level structures from word-occurrence patterns in text” (Chen, 2003, p. 139). Of particular importance in the context of this study is that co-word analysis is “a means to the elucidation of structures of ideas, problems, and so on, represented in appropriate sets of documents” (Whittaker, Courtial, & Law, 1989, p. 473). 

Although co-word analysis has its limitations, (e.g., Leydesdorff, 1997) primarily because of the
change of usage and meaning of words and the lack of context, such analysis has been considered particularly useful in tracking the development of scientific fields over time (Callon et al., 1991; Noyons & van Raan; Rip & Courtial, 1984), which represents another goal of this study.

Although citation analysis is not subject to the same limitation, it is a less direct indicator of cognitive structure. As already mentioned, studies using citations require post hoc assignment of topics. In addition, citation analysis of LIS is less effective in analyzing the cognitive structure of entire fields due to the different publication and citation practices of subfields, thus leaving even large subfields such as LS often invisible.

Selection of journals and articles. Articles from 16 LIS journals were chosen for inclusion in this study. The journals were selected from a ranked list of the most important journals in the field, according to deans and directors of American Library Association (ALA)-accredited, MLS programs in North America (Nisonger & Davis, 2005).

From this journal set, all research and review articles (10,344) published between 1988 and 2007 were included in the analysis.

Identification of the most frequently occurring LIS words and phrases. Word frequency is an important measure in content analysis. This measure is used to identify the most important research topics or concepts in a field by focusing on the most frequently occurring words.

In this study, we base all analyses on the 100 most frequently occurring LIS words or phrases. 

2014年8月14日 星期四

Huang, M. H., & Chang, Y. W. (2012). A comparative study of interdisciplinary changes between information science and library science. Scientometrics, 91(3), 789-803.

Huang, M. H., & Chang, Y. W. (2012). A comparative study of interdisciplinary changes between information science and library science. Scientometrics, 91(3), 789-803.

scientometrics

本研究利用圖書館學與資訊科學領域下各五種期刊於1978到2007年間論文引用的參考文獻,比較這兩個領域的跨學科(interdisciplinary)特性。跨學科性(interdisciplinarity)的定義為使用來自其他學科的知識(knowledge)、方法(methods)、技術(techniques)與設備(devices)成為科學活動的結果(Tijssen 1992),利用來自不同學科參考文獻的引用分布是經常採用的分析技術。研究結果顯示兩者的來源學科有很大不同:圖書館學的研究傾向於引用圖書資訊學(library and information science)、教育學(education)、企業/管理(business/management)、社會學(sociology)和心理學(psychology);然而資訊科學的研究引用大多來自圖書資訊學、一般科學(general science)、電腦科學(computer science)、科技(technology)和醫學(medicine)等學科。除了圖書資訊學本身以外,圖書館學引用的學科主要以社會科學為主,資訊科學的引用則主要來自於自然科學。

從引用比例的變化來看,圖書館學在引用圖書資訊學上有下降的趨勢,引用自教育學的比例則是上升,資訊科學來自電腦科學上的引用,其比例也是上升。

本研究以從Brillouin指標(Brillouin’s Index)測量兩個領域的跨學科性,Brillouin指標的計算方式如下:

N是觀察的數量(the number of observations),也就是參考文獻的總數,ni是屬於第i個類別的觀察的數量,也就是在第i個學科的參考文獻數量。從可以看到這兩個領域的跨學科性都逐年上升,並且資訊科學比圖書館學有較高的跨學科性。


Based on the research generated by five library science journals and five information science journals, library science researchers tend to cite publications from library and information science (LIS), education, business/management, sociology, and psychology, while researchers of information science tend to cite more publications from LIS, general science, computer science, technology, and medicine. This means that the disciplines with larger contributions to library science are almost entirely different from those contributing to information science.

However, a decreasing trend in the percentage of LIS in library science indicates that library science researchers tend to cite more publications from non-LIS disciplines. A rising trend in the proportion of references to education sources is reported for library science articles, while a rising trend in the proportion of references to computer science sources has been found for information science articles.

In addition, this study applies an interdisciplinary indicator, Brillouin’s Index, to measurement of the degree of interdisciplinarity. The results confirm that the trend toward interdisciplinarity in both information science and library science has risen over the years, although the degree of interdisciplinarity in information science is higher than that in library science.

The concept of interdisciplinarity has been discussed by many researchers (Huutoniemi et al. 2010; Leydesdorff and Probst 2009; Rosenfield 1992; Tijssen 1992), and can be defined as the use of knowledge, methods, techniques, and devices as a result of scientific activities from other fields (Tijssen 1992).

2013年12月19日 星期四

White, H. D. and McCain, K. W. (1998). Visualizing a discipline: An author co-citation analysis of Information Science, 1972–1995. Journal of the American Society for Information Science, 49, 327-355.

White, H. D. and McCain, K. W. (1998). Visualizing a discipline: An author co-citation analysis of Information Science, 1972–1995. Journal of the American Society for Information Science, 49, 327-355.

vis_paper

本論文探討作者共被引方法,並將其應用在資訊科學。這個研究分析了1972到1995年間12份資訊科學相關期刊內的作者共被引資料,以每八年為一期,所以整個24年研究共3期,每一期均找出被引用次數最多的前100位作者,整個期間共120位,其中的75位在三個時間都有出現。本研究使用的方法與結果分別如下
1) 對120位作者與其他作者的共被引次數形成的矩陣進行Pearson相關係數分析,再利用主成分分析(principal components analysis)與最大變異轉軸(varimax rotation)進行因素分析(factor analysis),了解資訊科學的專業(specialty)結構。以特徵值(eigenvalue)大於1決定抽取的因素數目,每一個因素代表一個專業,如果作者在某一特定的因素上具有0.3以上的負荷(loading),便視為引用者一般認為這位作者具有這個專業。由於作者可能在多個因素上都有超過0.3的負荷,因此每位作者可能會具有多種專業。在本研究中,共抽取出12個因素,可以解釋84%的變異情形,這些因素中前8個特徵值較大,可以從作者辨識的資訊科學專業為 a)設計與評估文件檢索系統的實驗檢索(experimental retrieval);b) 研究科學研究文獻關連的引用分析(citation analysis);c)應用於實際資料庫的實務檢索(practical retrieval);d) 從文字及書目資料分布規律探討數學模型的書目計量學(bibliometrics);e) 研究圖書館自動化、圖書館運作等議題的一般圖書館系統理論(general library systems theory);f) 研究資訊需求與使用的使用者理論(user theory);g) 研究科學的社會系統(social system of science)的科學傳播(scientific communication);h)OPAC ;另外幾個因素則由研究被引入資訊科學的其他領域學者組成。根據各專業上的作者交互情形以及下述映射圖的結果,資訊科學可以分為對於知識文獻以及其社會脈絡的分析研究和人-電腦-文獻的介面研究等兩個次學科。
2) 根據120位作者在3個時期的平均共被引次數,分析他們在各時期的代表性與影響力。
3) 以作者的共被引次數矩陣所產生的相關係數,也就是他們被引用者一般認定的相似性,做為他們之間的關連性,利用多維縮放技術ALSCAL,將每個時期前100位作者映射成圖形,使得共被引次數分布彼此相似的作者在產生圖形上的映射點有較近的距離。並以叢集分析技術CLUSTER進行完全連結叢集(complete linkage clustering),將作者根據他們之間的關連性分為次學科。結果發現,屬於同一個專業的作者在圖形上的映射點彼此間的距離比較近。並且如先前類似的研究所指出的,資訊科學很明顯地可以區分為資訊檢索及領域分析(domain analysis)等兩個次學科。比較不同時期的圖形,雖然少部分的作者映射點有明顯移動,但大多數的作者其映射點的位置相當穩定。
4) 從三個時期的映射圖上作者映射點位置的改變情形產生映射圖,表示作者引用形象(citation image)的改變。
5) 以經典作者(canonical auhtors)在三個時期的共被引相關係數為輸入,利用INDSCAL評估三個時期維度的重要性,從引用的角度驗證學科是否發生典範轉移的情形。結果發現表示「人-電腦-文獻」介面(human-computer-literatures interface)的第二個維度比起表示資訊科學主題專業的第一個維度在三個時期的重要性有大的變化,1972-1979年的第一時期這個維度的重要性不高,1980-1987年的第二時期其重要性則大幅增加,到了1988-1995年第三時期則稍微減少。許多研究者認為資訊科學在1980年代有典範轉移(paradigm shifting)發生,White and McCain上述的結果可以驗證這個現象。

We defined the authors of information science as all those cited in 12 journals, as listed below. Authors were ranked in order of citedness for the entire period covered by Social Scisearch, 1972–1995. Co-citation data were retrieved for all pairs in the top-ranked 120, from which we produced:
1) A factor analysis of the 120 authors for the entire 24-year span, 1972–1995, which reveals the specialty structure of the discipline. Factor analysis, unlike multi-dimensional scaling and clustering, can show an author’s contribution to more than one specialty.
2) Analyses of the 120 authors’ mean co-citation counts, which indicate their standing and influence in the discipline as of 1972–1979, 1980–1987, 1988–1995, and at the end of the three periods combined.
3) Two-dimensional maps of the top 100 authors in each of the 8-year periods (made with ALSCAL, the SPSS multidimensional scaling program) .
4) A map of authors whose ‘‘citation images’’ changed markedly over the years of our study.
5) A two-dimensional composite map of the authors who are in the top 100 in all three periods—some 75 in all. Their most cited works arguably make up the canonical literature of information science. Certain statistics generated by the mapping routine (INDSCAL, a part of ALSCAL) may bear on paradigm shift in the discipline.

In any field of scholarship, writers make judgments as to who has written on what, using what methods, and they reflect the judgments in their citing practices. Aggregated over time, these practices assume definite structure: Writers show commonalities in how they judge the subject matter, methodology, and intellectual style of other writers; for example, they often attach the same meanings and significance to precedent works (Cozzens, 1985; Small, 1978) .

It suggests how authors are commonly viewed on two dimensions, often interpretable as subject matter and style of work. ... Author clusters placed on these two dimensions can be interpreted as specialties within a discipline (White, 1990a, 1990b) .

What is actually mapped is an author’s citation image. Everyone ever cited has one, but only those who have been cited in many writings are likely to figure in ACA. In the latter case, the image has a constant part, the author’s identity as it is rendered in successive reference lists. The image also has a variable part, the gradually increasing set of other author-names that co-occur with a given author in those lists. At the end of a time period, ACA sums up the record by mapping the author as a single point among other selected author-points on the basis of the repeated co-occurrences. Authors with similar profiles of co-occurrences are displayed close together.

The decisive argument for ACA is that it enables one to see a literature-based counterpart of one’s own overview of a discipline.

As is well known, the closeness of author points on such maps is algorithmically related totheir similarity as perceived by citers. We use Pearson r as a measure of similarity between author pairs, because it registers the likeness in shape of their co-citation count profiles over all other authors in the set.

The raw co-citation counts were converted to Pearson r correlation matrices by the FACTOR routine in SPSS, and factors were extracted by principal components analysis with varimax rotation. The default criterion of ‘‘eigenvalues greater than one’’ determined the number of factors extracted.

The Pearson r correlation matrices for ALSCAL and CLUSTER in SPSS were generated with another SPSS rountine, CORRELATIONS ( cf. McCain, 1990) . They were treated as nonmetric (ordinal) similarity data in ALSCAL and grouped by the complete linkage method in CLUSTER. Subdisciplinary groupings of the author points on the maps are based on the dendograms from CLUSTER.

Authors in the top 100 in all three periods—‘‘the canonical 75’’—were separately mapped with INDSCAL, a routine in the ALSCAL bundle that does a specialized kind of multidimensional scaling. The input data to INDSCAL are judgments on the similarity of a set of stimuli by a set of judges. INDSCAL reveals not only the judges’ composite view of the stimuli in multidimensional space, but the weight each individual judge gives each dimension; INDSCAL is short for ‘‘individual differences scaling.’’ We used the individual weights in a new way to explore the notion of ‘‘paradigm shift’’ as it affects the canonical 75.

The two-dimensional space in which the authors appear is relative, not absolute, and it fails to capture certain relationships among oeuvres that appear in higher dimensionality.

Specialties
The results of the factor analysis, incorporating 24 years’ worth of data for the 120 authors, are presented in Table 3. ... Twelve factors were extracted; jointly (R2 ) , they explain 84% of the variance. ... The first eight factors alone explain 78% of the variance. All have seven or more authors with loadings greater than 0.60 and may be interpreted as specialties within the discipline.

The two biggest specialties, obviously, are experimental retrieval, which focuses on the design and evaluation of document retrieval systems, and citation analysis, which focuses on the interconnectedness of scientific and scholarly literatures, usually with data from ISI.

The third biggest specialty we have labeled practical retrieval. Unlike the experimental retrievalists, the authors in this group, rather than working with content-neutral indexing theory, thought experiments, or document testbeds, have tended to discuss retrieval in terms of ‘‘real world’’ databases; terms such as ‘‘INSPEC’’ or ‘‘DIALOG’’ occasionally profane their pens.

The next specialty we call bibliometrics—a word often used to subsume the specialty we labeled citation analysis. However, unlike the citationists, the authors who load primarily here, including the pioneers Lotka, Bradford, and Zipf, are most interested in mathematically modeling certain regularities in textual or bibliographic statistical distributions, irrespective of the literatures from which they come.

General library systems theory is a not altogether satisfactory name for a body of writings on library automation, library operations research, library and information service policy, retrieval system evaluation, and many other interconnected topics.

The specialty we call user theory is appropriately headed by Dervin, author of a highly cited chapter on ‘‘information needs and uses’’ in the 1986 ARIST. ... It will be seen that authors who write about literatures—the citationists, bibliometricians, and scientific communication people—never load above 0.30 on this factor, apparently because citers do not perceive their work as having the right psychological content. On the other hand, quite a few retrievalists load above 0.30, and this suggests the nature of the cognition involved. It has to do with problem-solving at the interface where literatures are winnowed down for users with: Question formulation, search strategies, information-seeking styles, relevance judgments, and the like.

Authors loading mainly on scientific communication all have strong disciplinary identities outside L&IS—for example, in sociology. They may be thought of as explicating the social systems of science, including those in which formal publication of results is an important (but not the only important) part. The sociologists among them all have loadings, some quite high, in citation analysis, confirming their relevance to the study of scientific literatures.

The design of computerized library catalogs, especially for subject searching, is the province of authors who load on OPACs (online public access catalogs) . It makes sense that leading authors here, such as Matthews, Hildreth, Cochrane, and Drabenstott, load secondarily in practical retrieval, just as several of the primary authors there, such as Borgman and Fidel, also turn up here.

As was said, the chief remaining factor seems a collection of authors in other disciplines from whom information science has imported ideas—e.g., cognitive science (Winograd) , information theory (Shannon) , computer science (Knuth)—that are all variously relevant to the central concern of information science, the human–computer–literature interface.

In fact, as both author cross-loadings and the maps below suggest, almost all of the factors or specialties in Table 3 can be aggregated upward into two larger subdisciplines: (1) The analytical study of learned literatures and their social contexts, comprising citation analysis and citation theory, bibliometrics, and communication in science and R&D; and (2) the study of the human–computer–literature interface, comprising experimental and practical retrieval, general library systems theory, user theory, OPACs, and indexing theory.

The Maps
Figures 2 through 4 are our 8-year period maps. We shall use them to explore the idea, introduced earlier, of two subdisciplines in information science.We operationalize this idea as the last two clusters joined in a complete-linkage clustering of 100 authors. These final clusters, which are brought together only after all closer ties have been exhausted, are separated by an angled line superimposed on each map.

We have not, as in the past, drawn lines around smaller clusters of authors corresponding to their specialties. The crowding of many names on the maps makes this difficult, and, besides, the specialties are better conveyed by the factor analysis of the earlier section. To a great extent, however, the authors forming specialties in the factor analysis will be found to have been placed near each other in the maps.

The first finding to note is the overall stability of information science, as here defined. Some author-points undergo remarkable changes of position from map to map, but many more authors stay put in discernible specialties. Fully 75, moreover, persist through all three maps.

We conclude that author co-citation analysis is useful for rendering the inertia of fields. In other words, it objectively captures the slow-changing divisions on which one’s subjective sense of ‘‘semi-permanent’’ disciplinary structure rests.
Co-citation analysis of papers, as opposed to authors, captures disciplinary history at a different, faster rate, which may better suit fields with livelier research fronts than information science.

However, ‘‘domain analysis,’’ as put forward by Hjørland and Albrechtsen (1995) , seems a more appropriate choice. It incorporates citation analysis and bibliometrics, but also a range of topics broader than what ‘‘bibliometrics’’ usually implies— for example, scholarly and professional communication, parts of sociology of science and sociology of knowledge, interdisciplinary linkages, discourse communities, and disciplinary vocabularies (cf. Beghtol, 1995) .
ACA’s confirmation of expert judgments by Hjørland and Albrechtsen, Persson, and the Vickerys is consistent with the claim that citation databases can be exploited for non-experts in a form of AI.

The axes in INDSCAL maps are not subject to rotation and are supposed to be maximally interpretable. Thus prompted, we think the horizontal axis conveys, as in past studies, the range of subject specialties within the subdisciplines of domain analysis and information retrieval. ... Coherent groups from left include the citationists, the arc of bibliometricians across the top and the philosophically orienting figures across the bottom, ‘‘generalist’’ writers such as Smith, Wilson, Saracevic, and Swanson, and the hard and soft retrievalists. The plot generally makes good sense. For example, it is easy to accept Bookstein, Tague-Sutcliffe, Kantor, Buckland, Vickery, and Shaw as transitional figures between the retrievalists and the bibliometricians.
The more interesting vertical axis reflects another subject-related continuum. Information science deals, we said earlier, with ‘‘the human–computer–literature interface.’’ If so, then the top pole represents a relative emphasis on literatures as objects of study, and the bottom, a relative emphasis on people or users. The same polarity can be inferred in earlier maps. Figure 4 showed that when a literature theoretician like Egghe enters, it is automatically at the top, whereas a user theoretician like Dervin is automatically placed at the bottom.

However, INDSCAL is expressly designed to reveal differences in the importance of each dimension to whoever is judging the similarity of stimuli. In our use of INDSCAL, the stimuli are the 75 authors, and the three periods are regarded as three separate ‘‘judges.’’
Usually, of course, persons are the judges in INDSCAL studies, and the ‘‘derived subject weights,’’ which are standard INDSCAL output, are taken to show the salience of each dimension to each person. In replacing individuals as judges with large numbers of citers, we are acting as if the citers collectively embodied the paradigm of information science in each 8-year period.
Accordingly, we interpret the derived subject weights for each period as indicating the relative importance of the dimensions within the paradigm. Thus, we can probe a hidden aspect of disciplinary history—whether key dimensions of the field were given about the same weight in all periods. If not, that would be consistent with a perception of paradigm shift.
Substantively, it is as if during 1972–1979 citers had regarded the range of specialties as by far the most important part of the information science paradigm, but then during 1980–1987 had taken much more cognizance of the differences in authors’ orientation toward literatures or users.

Perhaps the main weakness of this INDSCAL measure is that it is so indirect—that is, not clearly connected to specific papers with specific claims about the world. One expects evidence of paradigm shifts to leap from main texts, not references; from writers, not citers.
Though it might be used to discover paradigm shift, we think it has more promise as a means ofconfirming one. ... A shift detectable there implies not only that authors are promoting new lines of inquiry, but that citers are responding in such a way that the overall map of the discipline is changed.

Toward that account, ACA simultaneously provides both breadth and focus. It provides breadth by forcing contemplation of multiple specialties... It provides focus by forcing contemplation of particular authors, which is to say particular oeuvres and works. It also provides crude but unmistakable evidence of intellectual change.

The role of information science is to explicate the conceptual and methodological foundations on which existing systems are based’’ (Borko, 1968, p. 67). Or ‘‘Information science is the study of the means by which organised structures (which we call ‘information systems’) process recorded symbols to meet their defined objectives’’ (Hayes, 1985, p. 174) .
What they do study empirically, and uniquely, are problems associated with the human–literature barrier—the special difficulties of obtaining answers to questions from publications, in any medium, rather than persons. In other words, while many scholars seek to understand communication between persons, information scientists seek to understand communication between persons and certain valued surrogates for persons that literatures comprise (White, 1992).
This study requires a conceptual scheme that encompasses properties not only of literatures(e.g., size, growth rate, age, dispersion, authority levels, degree of summarization, quality of indexing) but also of people (e.g., interests and concerns, vocabularies, social ties, knowledge of existing systems, search styles, editorial strategies, resource environments).
The bond between domain analysts and retrievalists is their common interest in the literature barrier and related phenomena on both sides. The barrier in action is exemplified by information overload and underload—recurring topics for authors in both subdisciplines because they require both literatures and users to be discussed in a single framework, as implied by the second dimension of our maps.

2013年12月7日 星期六

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

本研究提出一個TTR-LDA-社群模型,這個模型以推論機制(inference mechanism)結合LDA(Latent Dirichlet Allocation)模型和Girvan-Newman社群偵測(community detection)演算法提供在網路資料上偵測社群並對這些社群進行主題探勘(topic mining)的功能,並且進而了解在社群上的主題隨時間推移的變化,處理的架構如下圖所示

本研究利用Delicious社會標籤系統(social tagging system)上從2005到2008年的資料進行研究。在社群偵測部分,首先建立網絡:根據使用者標籤的資源數量,選取前50000位標籤資源最多的使用者;然後對他們標籤的網頁進行統計,從其中選取10000個被最多使用者標籤的網頁。接著在上述的10000個網頁中,如果有這50000位使用者之間有任何兩位曾經標籤過相同的網頁,便在這兩位使用者之間產生一個連結。以50000個使用者為節點,同時以他們之間的連結為連結線,便可以建立一個共同書籤網絡(co-bookmark network)。並且為了研究網絡上社群結構的變化,並將整個期間的資料分為三個時段:分別為2005-2006、2007與2008年,相關的統計數據如下表:

本研究利用Girvan-Newman演算法找出標籤者(Tagger)社群,使得標籤者與同一社群內的其他標籤者比社群外的標籤者有較強的關係。這個演算法重複移去當時網絡上中介性(betweenness)最大的連結線,產生各種可能的網路劃分(network partition),測量每一種劃分下的群組性(modularity),也就是實際上社群內的成員彼此間的連結線數量與相同連結度的情況但隨機產生連結線的數量的差,群組性最大的劃分便是輸出結果。

另一方面,本研究利用TTR-LDA模型找出每個標籤者的主題分布以及主題內具有代表性的標籤,TTR-LDA模型修改自ACT(author-conference-topic)模式[11][12],是一個由標籤者做為第一層、標籤與資源為第三層、主題則為第二層,所構成的三層貝氏模型(three-layer Bayesian model)。

整合Girvan-Newman演算法和TTR-LDA模型的方法是以社群為單位,將社群內所有標籤者的主題分布進行平均做為該社群的主題分布,根據主題分布,選出機率值較大的主題做為社群的代表。比較不同時段社群共同的代表標籤衡量它們的相似性。

社群偵測的結果發現前五個最大的社群在四年裡占了絕大多數的比率,而且這個比率逐年增加。主題探勘的部分則測量TTR-LDA模型在不同主題數量上的複雜度(perplexity),發現150個主題時有最低的複雜度。因此,以下的研究便針對150個主題在前五個最大社群上的分布進行探討,計算它們的傳導性(conductance)與模組性。就社群模組性而言,最近一個時段(2008年)的結果比前三個時段還要高,其原因可能是因為經過一段時間後,社群的結構逐漸成熟,因此後期比前期更能產生較佳的社群。另外,將最後一個時段再細分為四個較小的時段則發現,較小時段的社群模組性比整年的結果來得高,本研究認為造成這種現象的原因可能是由於在不同的時段,大部分標籤者的書籤行為集中在不同的領域;當那些時段合併起來的時候,會展現標籤者在多個領域的興趣,使得社群內的叢集(clustering)特性較弱。

結果並可以發現前二十個主題都出現在不同時段的前五個最大社群裡,每個社群至少包括一個前十名的主題。此外,比較LDA、TTR-LDA和TTR-LDA-社群等三種模型在資源和標籤上的預測力,在回收率(recall rate)、精確率(precision)和F1指標上以TTR-LDA-社群為最佳。

In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism.

The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008.

Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time.

From a research perspective, these real-world networks display unique properties from the classical random graph model [3] in that most real word networks exhibit three common properties: the small-world property, power-law degree distribution and a high clustering coefficient or transitivity (indicating community structure) [7][8][9].

Thus, an important task in network analysis is to detect communities and explore their features, which can improve community-supporting services at the community-level in the context of a social tagging system.

Many studies in various disciplines have been devoted to community detection; however, few of them have systematically and quantitatively studied the profiles of those detected communities.

In this paper, we propose a TTR-LDA-Community model, which is an inferential combination of an extended LDA model and a betweenness-based community detection algorithm. It provides rich, systematic, and quantitative information about the profiles of detected communities.

In the context of social tagging systems, where multiple users are annotating resources, the resulting topics reflect a shared view of the document; and the tags of the topics reflect a common vocabulary.

Girvan and Newman extended the betweenness measure to edges and designed a clustering algorithm which gradually removes the edges with the highest betweenness value [4]. This algorithm has been improved through modularity; and the complexity is reduced from O(m2n) to O(mdlogn) where d is the depth of the dendrogram of the community structure [2].

Many studies provide various models and algorithms for topic mining and community detection; yet, few of them have integrated those models and algorithms, performed topic mining for detected communities, and analyzed how those identified topics change among communities over time.

The activity of social tagging consists of three major components: tag, tagger and resource. The experimental dataset contains all the triples of these three components and the time and date of their creation on Delicious from 2005 to 2008.

In data processing, all taggers were ranked by the number of resources they have bookmarked and the top 50,000 taggers were selected as the sample of taggers.

These taggers bookmarked a total of 354,522 web pages, which were sorted by the number of taggers who bookmarked them. The top 10,000 resources were selected as the sample of web pages, associated with which a dominant majority of tagging activities occurred.

Thus a co-bookmark network was built in which a connection between two users (within the sample of 50,000 taggers) is created if they bookmarked the same resources (within the sample of 10,000 web pages).

In addition, in order to observe the evolution of structure and motif of communities, the time span (2005-2008) was divided into three slices.


The model is illustrated in Figure 1. TTR-LDA is developed based on ACT model [11][12]. It is a three-layer Bayesian model with taggers tap in each post p as the first layer, tags t, and resource r as third layer and all the topics denoted as latent variable z as the middle layer.

The inference mechanism is used to infer the topic distribution over detected communities.

Each community includes a set of taggers, who have a stronger relationship with other taggers within the community than the taggers outside.

Based on the taggers’ information model, the probability distribution of each tagger over a set of topics is obtained by using the TTR-LDA model while the community structure of taggers is revealed by the community detection algorithm. The two sets of results are further integrated through an inference mechanism.




Results show that the number of users of the top five communities occupies a major proportion in the four years (2005-2008) and the proportion is increasing over time.

Perplexity is used to identify the number of topics [10], which arrives at the lowest point when the number of topics is 150. The interest model of each tagger in the top five largest communities is then built based on their topic distributions.

By using users’ interest models and the inference mechanism, a topic distribution of the largest community can be created (Figure 3). We can find that the topic distributions in a community are diverse because users’ relationships in that community are mainly based on their co-bookmark activities not the similarity of their interest model.

In order to observe the dynamic features of communities, we design an experiment as follows:
1) denote the five largest communities from each time slice in 2008 as community_i_t where t means the tth time slice in 2008 and i means the ith largest community in tth time slice;
2) compute the topic distribution for the five communities, which is stored as model_t_i_Topic(j), the
probability of jth topic in ith largest community in the tth time slice;
3) obtain the probability distribution of tags that are collected from all the posts generated during the specific time slice; the probability of one tag occurring in a topic shows the level of representativeness of the tag for that topic;
4) sort all the tags according to their probability value in each topic and select the 20 top ranked tags to represent the content of the topics; select the top 5 ranked topics to represent the theme of each community;
5) analyze the similarity between different communities from different time slice through computing how many tags are shared by the two different communities. More specifically, we compare current time slice with its previous time slice, for example, we compare community_i_t with community_j_t-1 (j=1, 2…5).

The size of communities along evolutionary lines fluctuates over time. For example, the size of the community about social networks in the 3rd time slice (community_1_3) is much larger (4,377) than that (521) in the 4th time slice (community_5_4).

Conductance (from multi-criterion scores) and modularity (from single criterion scores) are used to evaluate the quality of communities detected by the TTR-LDA-Community model [6].

The smaller the value of conductance is, the higher the granularity of a community is. Network community profile (NCP) is used to compute and display the value of conductance for communities [5].

Whiskers networks and rewired networks are adopted as two comparative aspects. Whiskers is defined as the maximal sub graphs that can be detached from the rest of the network by removing a single edge; and a rewired network is a random network that has the same nodes and the same degree distribution as the original network [5].

The conductance of communities of the rewired original network (blue line in the left figure), rewired random network (red dashed line in the left figure), the original whiskers network (blue line in the right figure), and the random whiskers network (red dashed line in the right figure) are calculated and shown in Figure 4.

In Figure 4, compared with the rewired network (left) and the rewired whiskers (right), 1) the original network displays a higher granularity of communities (a lower conductance value); 2) the value of conductance as the function of the size of communities in the original network and the original whiskers present a “V” shape, showing properties of a true large social networks [5]; 3) the original whiskers has the best community granularity (the lowest conductance) between size 10-100; and 4) the best community granularity of rewired original network is around 1000.

The modularity of communities in the four time slices of 2008 is better than that in 2005-2007. This is probably due to the fact that community structure grows mature gradually over time, creating better communities in later years than in earlier years.

Meanwhile, modularity of communities in the short-term (four sub periods in 2008) is larger than the long-term (2008). It can be explained that in different time periods, most taggers’ bookmarking activities are focused on different domains, so in a certain short-term time period, communities may be quite different from each other. However, when those time periods are merged together, the taggers show different interests in many domains; so the clustering feature within the communities becomes weaker.

Results show that the most popular topics are about bandslash fiction, fan fiction, and supernatural fiction (the top 3 popular topics). Communities with similar theme are ranked 3rd, 4th, and 5th in size; and the web resources with similar topics are ranked 500-600 of the top 1000 ranked resources in number of taggers associated with them.

The top 20 ranked topics in 1000 most popular resources can be found in 5 largest communities in different time periods. For each community, there exists at least one topic that is ranked top 10 in 1000 most popular resources (Table 3).

Topic distributions for each community are obtained respectively from LDA, TTR-LDA model, and TTR-LDACommunity model based on co-bookmark network in a given period (Oct. 2008–Dec. 2008). One resource and five tags are recommended for each post according to the results of three models separately.

The TTR-LDA and TTR-LDA-Community model show significant improvement for recommendation of tags and resources for post in terms of precision, recall and F1-Measure. TTR-LDA and TTR-LDA-Community have slightly improved performance for “tags for post”, while TTR-LDA-Community outperforms TTR-LDA on “resource for post”.