顯示具有 time analysis 標籤的文章。 顯示所有文章
顯示具有 time analysis 標籤的文章。 顯示所有文章

2015年3月30日 星期一

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

Chen, C., Ibekwe-SanJuan, F. and Hou, J. (2010), The structure and dynamics of cocitation clusters: A multiple-perspective cocitation analysis. Journal of the American Society for Information Science and Technology, 61 (7), 1386–1409. doi: 10.1002/asi.21309

確認科學領域的專業(specialties)本質是資訊科學的一項基本挑戰 (Morris & Van der Veer Martens, 2008; Tabah, 1999) 。由於1)可取用的書目資料來源愈來愈普及;2)網路上愈來愈多可提供分析與視覺化的電腦軟體工具;3)從多元來源而大量的資料吸收的要求愈來愈劇烈等原因,因此有愈來愈多的相關研究。共被引分析是對科學進行量化分析最常用的方法之一,特別是作者共被引分析 (author cocitation analysis, ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b)以及文件共被引分析 (document cocitation analysis, DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985)。作者共被引分析的目的在透過被相關文獻一起引用的作者群集,確認領域裡的專業。重要的作者共被引分析研究包括White & McCain (1998),這個研究以1972到1995年間12種資訊科學相關期刊的120位高被引作者進行作者共被引分析,研究結果發現當時的資訊科學分為兩個基本上彼此獨立的陣營:資訊檢索(information retrieval)與文獻(literature)。Zhao and Strotmann (2008a, 2008b) 以1996-2005年的資訊科學相關期刊資料重新進行了相同的研究,他們的結果發現了5個主要的專業:使用者研究(user studies)、引用分析(citation analysis)、實驗型檢索(experimental retrieval)、網路計量學 (Webometrics)以及知識領域的視覺化(visualization of knowledge domains),其中新興的兩個專業:網路計量學和知識領域的視覺化連繫了引用分析以及實驗型檢索,而使用者研究則是此時最大的專業。Aström (2007) 則是使用文件共被引分析的例子,他們分析了1990到2004年的21種圖書資訊學期刊,利用多維尺度法(multidimensional scaling, MDS)產生結果,他們的結果與White & McCain (1998)的研究類似,整個領域可分為兩個陣營,不過Aström (2007)的結果將稱為資訊尋求與檢索(information seeking and retrieval),而不是資訊檢索。

不管是作者共被引分析或是文件共被引分析其步驟大致如下:
1) 檢索引用資料。
2) 建構參考文件或作者共同被引用的矩陣。
3) 將共被引矩陣表示成節點與連結的圖(node-and-link graph)或是多維尺度法的組態(configuration),並且可以利用尋路網路(Pathfinder network scaling)或最小生成樹(minimum spanning tree)裁減連結。
4) 利用群集、社群發現(community finding)、因素分析(factor analysis)、主成分分析(principle component analysis)或者隱含語意索引(latent semantic indexing)等各種演算法確認專業。例如Morris & Van der Veer Martens (2008)、 Persson (1994)、 Tabah (1999)、 White & Griffith (1982)以及Janssens, Leta, Glänzel, and De Moor (2006)。
5) 根據群集成員間共同的主題(themes),解釋共被引群集的性質。通常需要豐富的領域知識,而且是一個花費大量時間與認知需求(cognitively demanding)的工作。

本研究對於作者共被引以及文件共被引形成的群集進行結構與動態的描述與解釋,分析的資料為1996到2008年間的12種資訊科學(information science)領域相關期刊,共計10853筆書目紀錄,引用的參考文獻為129060筆,引用次數為206180,而參考文獻的作者共有58711位。本研究以餘弦(cosine)測量作者或文件之間的關連大小,做為節點間的連結,建立網路;然後計算從原先網路導出的Laplacian矩陣(Laplacian matrices)的特徵向量(eigenvectors)找出群集。這種利用標準線性代數的頻譜群集(spectral cluster)演算法,較其他的群集演算法更有效率,而且因為不需要假設群集的形式,所以更有彈性與強健。標註群集方面則是利用引用文獻論文的詞語與摘要句,詞語包括題名與摘要中出現的名詞片語與索引詞(index terms),利用 tf*idf (Salton, Yang, & Wong, 1975)、對數似然比(log-likelihood ratio, LLR)測試 (Dunning, 1993)以及相互資訊(mutual information, MI)等三種資訊做為判斷的參考。摘要句則是從題名與摘要尋找最有代表性的句子,例如以Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007)對句子進行排序。




A multiple-perspective cocitation analysis method is introduced for characterizing and interpreting the structure and dynamics of cocitation clusters.

The generic method is applied to a three-part analysis of the field of information science as defined by 12 journals published between 1996 and 2008: (a) a comparative author cocitation analysis (ACA), (b) a progressive ACA of a time series of cocitation networks, and (c) a progressive document cocitation analysis (DCA).

Identifying the nature of specialties in a scientific field is a fundamental challenge for information science (Morris & Van der Veer Martens, 2008; Tabah, 1999).

The growing interest in mapping and visualizing the structure and dynamics of specialties is because of a number of reasons:
1. Widely accessible bibliographic data sources such as the Web of Science, Scopus, and Google Scholar (Bar-Ilan, 2008; Meho & Yang,2007) as well as domain-specific repositories such as ADS (http://www.adsabs.harvard.edu/) and arXiv (http://arxiv.org/).
2. Freely available computer programs and Web-based general-purpose visualization and analysis tools such as ManyEyes (http://manyeyes.alphaworks.ibm.com/) and Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/; Batagelj & Mrvar, 1998), special-purpose citation analysis tools such as CiteSpace (http://cluster.cis.drexel.edu/&u0007E;cchen/citespace/; Chen, 2004; Chen, 2006), and social network analysis such as UCINET (http://www.analytictech.com/ucinet6/ucinet.htm).
3. Intensified challenges for digesting the vast volume of data from multiple sources (e.g., e-Science, Digging into Data (http://www.diggingintodata.org/), cyber-enabled discovery, SciSIP; Lane, 2009).

Cocitation studies are among the most commonly used methods in quantitative studies of science, especially including author cocitation analysis (ACA; Chen, 1999; Leydesdorff, 2005; White & McCain, 1998; Zhao & Strotmann, 2008b) and document cocitation analysis (DCA; Chen, 2004; Chen, 2006; Chen, Song, Yuan, & Zhang, 2008; Small & Greenlee, 1986; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985).

For instance, once cocitation clusters are identified, assigning the most meaningful labels for these clusters is currently a challenging task because any representative labels of clusters must characterize not only what clusters appear to represent, but also the salient and unique reasons for their formation.

The new procedure reduces analysts' cognitive burden by automatically characterizing the nature of a cocitation cluster in terms of (a) salient noun phrases extracted from titles, abstracts, and index terms of citing articles and (b) representative sentences as summarizations of clusters.

ACA aims to identify underlying specialties in a field in terms of groups of authors who were cited together in relevant literature.

White & McCain (1998) presented a comprehensive view of information science based on 12 journals in library and information science across a 24-year span (1972–1995). It analyzed cocitation patterns of 120 most-cited authors with factor analysis and multidimensional scaling. The authors drew upon their extensive knowledge of the field and offered an insightful interpretation of 12 specialties identified in terms of 12 factors. The most well-known finding of the study is that information science at the time consisted of two essentially independent camps, namely, the information retrieval camp and the literature camp, including citation analysis, bibliometrics, and scientometrics.

Zhao and Strotmann (2008a, 2008b) followed up White and McCain's study using the same set of 12 journals and the same number of 120 cited authors in an updated time frame of 1996-2005. ... Zhao and Strotmann (2008b) found five major specialties and manually labeled them as user studies, citation analysis, experimental retrieval, Webometrics, and visualization of knowledge domains. In contrast to the findings of (White & McCain, 1998), experimental retrieval and citation analysis retained their fundamental roles in the field, and the user studies specialty became the largest specialty. Webometrics and visualization of knowledge domains appeared to make connections between the retrieval camp and the citation analysis camp.

A DCA by Aström (2007) studied papers published between 1990 and 2004 in 21 library and information science journals. Results were depicted in multidimensional scaling (MDS) maps. Aström's study also identified the two-camp structure found by (White & McCain, 1998). On the other hand, Aström found an information seeking and retrieval camp, instead of the information retrieval camp as in (White and McCain).

Although manually labeling a cocitation cluster can be a very rewarding process of learning about the underlying specialty and result in insightful and easy to understand labels, it requires a substantial level of domain knowledge and it tends to be time-consuming and cognitively demanding because of the synthetic work required over a diverse range of individual publications.

Traditionally, researchers often identify the nature of a cocitation cluster based on common themes among its members. ... The emphasis on common areas is a practical strategy; otherwise, comprehensively identifying the nature of a specialty can be too complex to handle manually.

Many researchers have studied the structural and dynamic properties of specialties in information science in terms of clusters, multivariate factors, and principle components (Morris & Van der Veer Martens, 2008; Persson, 1994; Tabah, 1999; White & Griffith, 1982).

A recent study of information science (Ibekwe-SanJuan, 2009) mapped the structure of information science at the term level using a text analysis system TermWatch and a network visualization system Pajek, but it did not address structural patterns of cited references.

Researchers also studied the structure of information science qualitatively, especially with direct inputs from domain experts. For example, Zins conducted a Critical Delphi study of information science, involving 57 leading information scientists from 16 countries (Zins, 2007a, 2007b, 2007c, 2007d).

Janssens, Leta, Glänzel, and De Moor (2006) studied the full-text of 938 publications in five library and information science journals with latent semantic analysis (LSA; Deerwester, Dumais, Landauer, Furnas, & Harshman, 1990) and agglomerative clustering. They found an optimal 6-cluster solution in terms of a local maximum of the mean silhouette coefficients (Rousseeuw, 1987) and a stability diagram (Ben-Hur, Elisseeff, & Guyon, 2002). Their clusters were labeled with single-word terms selected by tf*idf (p. 1625), which are not as informative as multiword terms for cluster labels.

Klavans, Persson, and Boyack (2009) recently raised the question of the true number of specialties in information science. They suspected that the number is much more than the 11 or 12 as reported in ACA studies such as (White & McCain, 1998) and (Zhao & Strotmann, 2008a, 2008b), but significantly fewer than the 72 reported in their own study, which is also based on the 12 journals between 2001 and 2005.

The 12-journal Information Science dataset, retrieved from the Web of Science, contains 10,853 unique bibliographic records, written by 8,408 unique authors from 6,553 institutions and 89 countries. These articles cited 129,060 unique references for a total of 206,180 times. They cited 58,711 unique authors and 58,796 unique sources.

The traditional procedure of cocitation analysis for both DCA and ACA comprises the following steps:
1. Retrieve citation data from sources such as the Science Citation Index (SCI), Social Science Citation Index (SSCI), Scopus, and Google Scholar.
2. Construct a matrix of cocited references (DCA) or authors (ACA).
3. Represent the cocitation matrix as a node-and-link graph or as a multidimensional scaling (MDS) configuration with possible link pruning using Pathfinder network scaling or minimum spanning tree algorithms.
4. Identify specialties in terms of cocitation clusters, multivariate factors, principle components, or dimensions of a latent semantic space using a variety of algorithms for clustering, community finding, factor analysis, principle component analysis, or latent semantic indexing.
5. Interpret the nature of cocitation clusters.

The interpretation step is the weakest link. It is time-consuming and cognitively demanding, requiring a substantial level of domain knowledge and synthesizing skills. In addition, much of attention routinely focuses on cocitation clusters per se, but the role of citing articles that are responsible for the formation of such cocitation clusters may not be always investigated as an integral part of a specialty.

Our new method extends and enhances traditional cocitation methods in two ways: (a) by integrating structural and content analysis components sequentially into the new procedure and (b) by facilitating analytic tasks and interpretation with automatic cluster labeling and summarization functions. The new procedure is highlighted in yellow in Figure 2, including clustering, automatic labeling, summarization, and latent semantic models of the citing space (Deerwester et al., 1990).

Our new procedure adopts several structural and temporal metrics of cocitation networks and subsequently generated clusters.

Structural metrics include betweenness centrality, modularity, and silhouette.

Temporal and hybrid metrics include citation burstness and novelty

The betweenness centrality metric is defined for each node in a network. It measure the extent to which the node is in the middle of a path that connects other nodes in the network (Brandes, 2001; Freeman, 1977). High betweenness centrality values identify potentially revolutionary scientific publications (Chen, 2005) as well as gatekeepers in social networks.

In the context of this study, the modularity Q measures the extent to which a network can be divided into independent blocks, i.e., modules (Newman, 2006; Shibata, Kajikawa, Taked, & Matsushima, 2008).

The silhouette metric (Rousseeuw, 1987) is useful in estimating the uncertainty involved in identifying the nature of a cluster.

Burst detection determines whether a given frequency function has statistically significant fluctuations during a short time interval within the overall time period.

Sigma is introduced in (Chen, et al., 2009a) as a measure of scientific novelty. ... In this study, Sigma is defined as (centrality + 1)burstness such that the brokerage mechanism plays more prominent role than the rate of recognition by peers.

We adopt a hard clustering approach such that a cocitation network is partitioned to a number of nonoverlapping clusters.

In this article, cocitation similarities between items i and j are measured in terms of cosine coefficients.

A good partition of a network would group strongly connected nodes together and assign loosely connected ones to different clusters. This idea can be formulated as an optimization problem in terms of a cut function defined over a partition of a network. Technical details are given in relevant literature (Luxburg, 2006; Ng, Jordan, & Weiss, 2002; Shi & Malik, 2000).

Spectral clustering is an efficient and generic clustering method (Luxburg, 2006; Ng et al., 2002; Shi & Malik, 2000). It has roots in spectral graph theory. Spectral clustering algorithms identify clusters based on eigenvectors of Laplacian matrices derived from the original network.

Spectral clustering has several desirable features compared to traditional algorithms such as k-means and single linkage (Luxburg, 2006):
 • It is more flexible and robust because it does not make any assumptions on the forms of the clusters,
• it makes use of standard linear algebra methods to solve clustering problems, and
• it is often more efficient than traditional clustering algorithms.

Candidates of cluster labels are selected from noun phrases and index terms of citing articles of each cluster. These term are ranked by three different algorithms. In particular, noun phrases are extracted from titles and abstracts of citing articles. The three term ranking algorithms are tf*idf (Salton, Yang, & Wong, 1975), log-likelihood ratio (LLR) tests (Dunning, 1993), and mutual information (MI).

Each cocitation cluster is summarized by a list of sentences selected from the abstracts of articles that cite at least one member of the cluster.

In this study, sentences are ranked by Enertex (Fernandez, SanJuan, & Torres-Moreno, 2007). Given a set S of N sentences, let M be the square matrix that for each pair of sentences gives the number of nominal words in common (nouns and adjectives).

In this study, summarization sentences were also ranked by two new functions gtf and gftidf , which are further simplified approximations of the energy function E.

The ACA and DCA studies described in this article were conducted using the CiteSpace system (Chen, 2004; Chen, 2006). CiteSpace is a freely available Java application for visualizing and analyzing emerging trends and changes in scientific literature.

CiteSpace supports a unique type of cocitation network analysis—progressive network analysis—based on a time slicing strategy and then synthesizing a series of individual network snapshots defined on consecutive time slices. Progressive network analysis particularly focuses on nodes that play critical roles in the evolution of a network over time. Such critical nodes are candidates of intellectual turning points.

In summary, (a) spectral clustering and factor analysis identified about the same number of specialties, but they appeared to reveal different aspects of cocitation structures and (b) cluster labels chosen from citers of a cluster tend to be more specific terms than those chosen by human experts.

We found the comparison with the study of Zhao and Strotmann very valuable. It offered us an opportunity to compare the analysis conducted by human experts to the interpretation cues provided by our automatic labeling and summarization methods.

Spectral clustering for the purpose of network decomposition is exclusive in nature although in reality it is often sensible to allow overlapping clusters because of multiple roles individual entities may play.

Spectral clustering of cocitation networks tends to generate distinct clusters with high precision, whereas human experts tend to aggregate entities into broadly defined clusters.

In conclusion, the new cocitation analysis procedure has the following advantages over the traditional one:
• It can be consistently used for both DCA and ACA.
• It uses more flexible and efficient spectral clustering to identify cocitation clusters.
• It characterizes clusters with candidate labels selected by multiple ranking algorithms from the citers of these clusters and reveals the nature of a cluster in terms of how it has been cited.
• It provides metrics such as modularity and silhouette as quality indicators of clustering to aid interpretation tasks.
• It provides integrated and interactive visualizations for exploratory analysis.

Modularity and silhouette metrics provide useful quality indicators of clustering and network decomposition.

2015年3月24日 星期二

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

由於發現新的物種、疾病與社交模式,產生新的研究主題與專業 (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012),經過一段時間後,相關的研究社群會成長或是規模改變,有些主題仍然持續,但有些則是消失 (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010)。已有許多研究利用書目資料來確認研究的專業,例如Kessler (1963)的論文書目耦合網路(bibliographic coupling networks)、Small (1973)的論文共被引網路(paper co-citation networks)、White 與 McCain (1998)的作者共被引網路(author co-citation networks)以及White (2003)的尋路網路 (pathfinder networks),Callon、Courtial 與 Laville (1991)、Ding、Chowdhury 與 Foo (2000)、Milojevic、Sugimoto、Yan 與 Ding (2011)則是使用詞語共現網路 (co-word networks)。這些研究各自在不同研究層次確認研究主題,例如論文層次有 Chen (2004, 2006)、 Kessler (1963)和 Small (1973),作者層次有 Clauset, Newman, & Moore (2004)、White & McCain (1998)和 White (2003),期刊層次如 Glänzel & Schubert (2003)、 Leydesdorff & Vaughan (2006),以及領域層次有 Janssens, Zhang, Moor, & Glänzel (2009)、Rafols & Leydesdorff (2009)、Zhang, Liu, Janssens, Liang, & Glänzel (2010)。較低的研究實體層級,如論文與作者,研究可以從領域內發現其他的主題或專業;但在期刊或領域等較高的層次,通常從更完整的資料中確認出次領域。確認主題的方法則有因素分析(factor analysis)和多維尺度(multidimensional scaling)等傳統的群集技術以及連結線中心性(edge betweenness)、群組性(modularity)和混合群集(hybrid clustering)等較新技術的應用。本研究(Yan, 2014)則是利用主題模型(topic model)確認研究主題,並提出主題延續性(topic continuity)及主題普遍性(topic popularity)等兩項動態特性來分析研究主題。應用主題模型技術考察主題動態的方法,包括事後分析(post hoc analysis)(例如: Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008)、分段法(segmented approaches) (例如:Bolelli, Ertekin,Zhou, & Giles, 2009)以及連續時間模型(continuous-time model) (Wang & McCallum, 2006)等。本研究採用事後分析,利用文件中各主題的機率分布評估主題存在的機率。

本研究針對每一年分別產生一個主題模型,對於每一個主題找出後一年最有可能的主題,評估兩個主題相似的方式是利用改良自Kullback-Leibler差異 (Kullback-Leibler divergence, KLD)的Jensen–Shannon差異 (Jensen–Shannon divergence, JSD),兩個模型P和Q的JSD計算方式為JSD(P||Q) = 1/2KLD(P||M)+1/2KLD(Q||M),KLD是兩個模型Kullback-Leibler差異,M=1/2(P+Q)。每一個主題後一年最有可能的主題是擁有最小JSD的主題,其分數JSD稱為JJSDS,運用JJSDS的變化趨勢計算主題的連續性,然後以z score進行標準化。
評估各主題的普遍性則是計算它們在該年度文件上平均的機率值,並以z score進行標準化,愈大的機率值表示該主題在當年度愈普遍,然後分析主題普遍性的變化趨勢,。

本論文的研究資料為2001到2011年的圖書資訊學(library and information science)出版品,包括期刊論文、書評及研討會論文等,採用論文的題名做為分析資料,共27,796 篇論文。每一年的主題數目都設為20。結果顯示在網路資訊檢索(web information retrieval)、引用及書目計量學(citation and bibliometrics)、系統及技術(system and technology)、健康科學(health science)等主題有較高的平均普遍性;h指標(h-index)、線上社群(online communities)、資料保存(data preservation)、社群媒體(social media)和網站分析(web analysis)等則是圖書資訊學裡愈來愈普遍的主題。研究結果的主題與過去的研究相符合,但這篇論文的貢獻在於對於研究主題的動態進行分析。

Dynamic development is an intrinsic characteristic of research topics. To study this, this paper proposes two sets of topic attributes to examine topic dynamic characteristics: topic continuity and topic popularity.

Topic continuity comprises six attributes: steady, concentrating, diluting, sporadic, transforming, and emerging topics; topic popularity comprises three attributes: rising, declining, and fluctuating topics.

These attributes are applied to a data set on library and information science publications during the past 11 years (2001–2011).

Results show that topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity; topics on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular in library and information science.

Dynamics is a constant theme in scientific explorations. Research communities may grow or change in size; new species, diseases, or societal patterns may be discovered; and new research topics and specialties may be introduced (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012). Over time, some topics are continuously investigated while others appear or disappear (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010). Therefore, it is of great importance to examine research dynamics to understand the evolving cognitive structures of research domains.

Pioneering studies of paper bibliographic coupling networks (Kessler, 1963), paper co-citation networks (Small, 1973), author co-citation networks (White & McCain, 1998), pathfinder networks (White, 2003) and co-word networks (e.g., Callon, Courtial, & Laville, 1991; Ding, Chowdhury, & Foo, 2000; Milojevic, ´ Sugimoto, Yan, & Ding, 2011) were capable of identifying research specialties from bibliographic data effectively.

However, findings from these studies remained largely static and thus only yielded fixed perspectives on the cognitive structure of research domains.

To examine research dynamics,this study uses a topic modeling technique and proposes two sets of topic attributes–topic continuity and topic popularity.

• How to use topic modeling techniques to study research dynamics?
• What quantitative measurements can be used to describe topic dynamics?
• What topics are present in library and information science? What are their dynamic characteristics?

This subsection reviews the network-based approaches of identifying research topics and specialties. These approaches have been applied to several research levels, including the paper-level (e.g., Chen, 2004, 2006; Kessler, 1963; Small, 1973), the author-level (e.g., Clauset, Newman, & Moore, 2004; White & McCain, 1998; White, 2003), the journal-level (e.g., Glänzel & Schubert, 2003; Leydesdorff & Vaughan, 2006), and the field-level (e.g., Janssens, Zhang, Moor, & Glänzel, 2009; Rafols & Leydesdorff, 2009; Zhang, Liu, Janssens, Liang, & Glänzel, 2010).

Most above-mentioned work used co-occurrence networks as the research instrument.

Analyses on lower level research entities, such as papers and authors, usually identified topics and specialties from small but well-defined research fields; whereas analyses on higher level research entities, such as journals and fields, attempted to identify subfields and subdomains from more comprehensive data sets.

Both classic clustering techniques (e.g., factor analysis and multidimensional scaling) as well as modern techniques (e.g., edge betweenness, modularity, and hybrid clustering) have been applied.

Recently, studies have attempted to add dynamic analyses by utilizing multiple time intervals.

Several approaches on slicing time intervals are available: intervals that have the same amount of references (e.g., Radicchi et al., 2009), intervals that have the same number of publications (e.g., Sugimoto, Li, Russell, Finlay, & Ding, 2011; Yan & Sugimoto, 2011), same-length intervals (e.g., Åström, 2007; Milojevic´ et al., 2011), and accumulative intervals (e.g., Barabási et al., 2002; Yan & Ding, 2009).

These studies laid valuable methodological basis for dynamic analyses of cognitive structures of research fields; however, networks of different time frames were largely analyzed distinctively and a more integrated examination was lacking.

In the meantime, empirically, network-based clustering results may require domain expertise to effectively interpret obtained results.

Topic modeling techniques use probabilistic models to assign papers, journals, or authors to clusters. A topic can be defined as a probability distribution over terms in a vocabulary (Blei & Lafferty, 2007). Latent Dirichlet Allocation (LDA) model, a classic topic model, was proposed by Blei et al. (2003). The model predicates that words for each paper are derived from a mixture of topics and each topic follows a multinomial distribution.

One recent update of the LDA model is the supervised LDA model. It makes the analyses of multi-labeled corpora (e.g., tags from delicious.com and various classifications) possible. Blei and McAuliffe’s (2010) version of supervised LDA can successfully address this challenge, but a document can only be assigned with one label.

Ramage, Hall, Nallapati, and Manning (2009) offered an approach which enabled the multi-label assignment. Their supervised labeled LDA (L-LDA) associated one label with one topic and allowed the model to learn word-label relations.

Through topic modeling techniques, topic dynamics has been examined mainly through the following approaches: post hoc analysis (e.g., Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008), segmented approaches (e.g., Bolelli, Ertekin,Zhou, & Giles, 2009), and continuous-time model (Wang & McCallum, 2006).

Post hoc analysis uses topic-document probability distributions to evaluate the presence of identified topics.

Segmented approaches build the dynamic component in the probabilistic model. It assumes that the state of topics at a single time point is independent from all other time points and divides document corpora into segments that have contingent time stamps (Bolelli et al., 2009).

Continuous-time model is a non-Markov model proposed by Wang and McCallum (2006), where they found the non-Markov model provides better prediction and more interpretable topical trends.

In this study, a post hoc dynamic analysis using the ACT model is selected because of its marked performance (Tang et al., 2008) as well as its advanced input and output support.

Topic dynamics is calculated through the Author-Conference-Topic (ACT) model (Tang et al., 2008).

Specifically, i is the topic distribution for document i. Mean ( ¯), therefore, is a direct quantitative measurement to assess topic popularity: the higher the ¯, the more visible the topic, and thus the more popular that topic is (Griffiths & Steyvers, 2004).

Because the data set spans 11 years, 11 independent ACT models were run, one for each year of the data set based on year of publication.

The Jensen–Shannon divergence (JSD) was used as the similarity measurement to quantify the topic similarity between different word-topic distributions. ... JSD is a symmetrized and smoothed version of the Kullback–Leibler divergence (KLD). ... As a divergence measure, the smaller the JSD, the higher the similarity is.

In order to track the same topic from two adjacent time intervals, the minimum value for each row of a JSD matrix was used, referred to as the joint JSD score (JJSDS): MIN(JSD Matrix(i,j)), for j = 1:n. ... Applying the same approach to each pair of adjacent time slices, for each topic, an array of JJSDS can be obtained.

The attributes of steady, concentrating, and diluting topics focus on the overall topical characteristics whereas the attributes of sporadic, transforming, and emerging topics focus on the topical characteristics of a specified time frame. Therefore, these attributes are not mutual exclusive, suggesting that a topic can be a concentrating topic overall, and in the meantime, related topics were added and thus qualifying it for a transforming topic.

The data set contains publications of all journals indexed in the 2011 version of the Journal Citation Report in the Information Science & Library Science subject category. Articles, proceeding papers, and review articles published within these journals from 2001 to 2011 were downloaded for analysis (downloading time: October 2012). Stop words were then removed from publications’ titles. Publications without titles, authors, or journal names were removed from the data set. The final data set comprised 27,796 papers.

The number of topics is set at 20: this number considers the size of the paper corpus as well as previous empirical studies on the cognitive structure of library and information science (e.g., Milojevic´ et al., 2011; Sugimoto et al., 2011; White & McCain, 1998; Zhao & Strotmann, 2008). For reasons of consistency, the same number of topics was identified for each year of the data set.

In this subsection, we first present histograms made from values in Jensen–Shannon divergence (JSD) matrices (Fig. 4). These histograms provide a direct perception on how research topics in library and information science are related as measured by JSD. This subsection then introduces all 20 topics in each year from 2001 to 2011 as well as how topic continuity and popularity attributes are applied to these topics (Fig. 5).

Fig. 4 uses histograms to visualize JSD values for each pair of adjacent years. Because there are 20 topics for each year, the number of data points in each histogram is 400 (20 × 20). This number is 4000 for the histogram in the lower right section of Fig. 4, as it uses JSD values for all pairs of adjacent years.

This study finds that in library and information science, research topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity over the past decade (from 2001 to 2011).

Research on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular topics.

Overall, findings of this study are consistent with previous studies using co-word, co-citation, and topic modeling techniques.

For instance, a co-word study by Milojevic´ and colleagues (2011) has found that title terms “citation”, “impact factor”, and “web” have a rising usage from 1989 to 2008.

Other related dynamic studies that cover the target time frame of the current study (2001–2011) include Åström’s (2007) study on examining library and information science research front, where the study found that webometrics and information-seeking and retrieval have become dominating research areas between 2000 and 2004.

This finding has been verified by Klavans and Boyack (2011) where the authors used the global map (i.e., the map of science) to enhance to accuracy of local maps (i.e., the contextual map of information science). They identified five core areas in information science, including information-seeking behavior, computer-enhanced retrieval, scientometrics, co-citation analysis, and citation behavior.

Besides the contextual analysis of information science, structural analysis has also been achieved from a time-series empowered author co-citation and document co-citation analysis (Chen, Ibekwe-SanJuan, & Hou, 2010). Through the application of a series of structural metrics such as centrality measures, modularity and silhouette, a clear cognitive structure of information science was attained in that the research areas of interactive information retrieval, academic web, information retrieval, citation behavior, and h-index have gained a particular popularity from 1996 to 2008.

In addition to journal publications, Sugimoto and colleagues (2011) applied a LDA model to library and information science dissertations and demonstrated dissertations as an important communicative genre. Their study indicated that between 2000 and 2009, internet and information retrieval related topics were the central dissertation research themes.

The contribution of the current study is that it proposes two sets of quantitative topic attributes. These attributes have streamlined the dynamic analysis of research topics and specialties and have further complemented co-occurrence-based studies.

This paper has identified dynamic characteristics of topics in library and information science; however, limited information can be told about the mechanisms that resulted in such characteristics. That being said, the study is unable to pinpoint, for instance, whether the growing popularity of network and citation studies is the result of a growing research community, a drive by the commercial market, a stimulus from funding agencies, or a combination of these or other unlisted factors.

Popular topics may be associated with research communities that are expanding in size and/or tend to have higher productivity. Conversely, less popular topics may be associated with communities that are shrinking and/or have a reduced productivity. Topic continuity and popularity attributes reflect research specialties’ development in scientific communities, which is further guided by science policies and the attention of the general public.

In informetrics, studies have mainly focused on analyzing the performance and the social and cognitive implications of several types of research entities, including papers, authors, institutions, journals, and fields. Authors and institutions are typically used to examine social relations in academia; while journals and fields are predominantly used to investigate the cognitive structure of research domains.

Topic analysis can precisely provide a more refined assessment by clustering research papers based on certain probability distributions. Because of such quantitative results, a more integrated dynamic cognitive analysis is thus possible, as exemplified through the current study.

Topic analysis will be further developed by overlaying topics with author communities to explore the interwoven relationships between research topics and research communities (e.g., Yan et al., 2012); by overlaying topics with funding data to investigate the “lead-lag” relationship between funding support and productivity (e.g., Shi et al., 2010); by applying topic models to different genres to study research immediacy (e.g., Ding et al., 2013); and by overlaying topics with citation data to examine the relationships between topics and impact.

2014年2月28日 星期五

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of American Society for Information Science and Technology, 57(3), 359-377.

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature.  Journal of American Society for Information Science and Technology, 57(3), 359-377.

information visualization

本研究提出一個整合研究專業(specialty)的研究前沿(research front)以及其引用的知識基礎(intellectual base)的視覺化介面。本論文定義研究前沿為研究專業上一組急遽出現的概念(concepts)與研究議題(research issues);研究前沿的知識基礎則是包含這些概念與研究議題的論文引用或者共同被引用的論文。在針對某一個專業進行其研究前沿與知識基礎進行視覺化時,首先蒐集專業相關的論文,從這些論文抽取代表研究前沿的詞語,並以論文所引用或共被引的論文做為專業的知識基礎,建立分別代表研究前沿的詞語和知識基礎的論文的二方網路(bipartite networks)以同時呈現研究前沿的相關概念與研究議題以及知識基礎的論文。在建立起來的網路上透過詞語和論文形成的叢集可以發現重要的研究前沿和知識基礎,藉由詞語呈現叢集的概念與研究議題更能有效地表達研究前沿的意涵,並且如果加上論文的發表時間來分析,可以從急遽出現在較多論文的相關詞語找出發展中的研究前沿。此外,對於網路進行中介中心性(centrality of betweenness)分析可以發現研究前沿間具有樞紐地位的論文,並且透過Pathfinder演算法可以發現論文間的主要關連。
A specialty is conceptualized and visualized as a time-variant duality between two fundamental concepts in information science: research fronts and intellectual bases.
A research front is defined as an emergent and transient grouping of concepts and underlying research issues.
The intellectual base of a research front is its citation and co-citation footprint in scientific literature— an evolving network of scientific publications cited by
research-front concepts.
The concept of a research front was originally introduced by Price (1965) to characterize the transient nature of a research field. Price observed what he called the immediacy factor: There seems to be a tendency for scientists to cite the most recently published articles. In a given field, a research front refers to the body of articles that scientists actively cite.
A specialty can be conceptualized as a time-variant mapping from its research front to its intellectual base.
Typical questions regarding a research front may include:
How did it get started? What is the state of the art? What are the critical paths in its evolution?
To address such questions, we need to detect and analyze emerging trends and abrupt changes associated with a research front over time. We also need to identify the focus of a research front at a particular time in the context of its intellectual base, to reveal significant intellectual turning points as a research front evolves, and to discover the interconnections between different research fronts.
Braam, Moed, and Raan (1991) defined a specialty as “focused attention by a number of scientific researchers to a set of related research problems and concepts” (p. 252). They studied the continuity and stability of a specialty in terms of the similarity between co-citation clusters across consecutive years. The similarity between two co-citation clusters is determined by comparing aggregated word profiles of the clusters.
In part, this is because we define a research front differently to emphasize emerging trends and abrupt changes as the defining features of a research front. A research front is the domain of a time-variant mapping, and its intellectual base is the co-domain of the mapping.
Griffith et al. (1974) found that between-cluster co-citation links tend to be weaker than within-cluster co-citation links. ... To understand how specialties and different thematic trends interact with each other, it is essential to study the nature of long-range, between-cluster links and understand why articles in different specialties were connected.
Labeling clusters is concerned with the clarity and interpretability of co-citation clusters. The standard approach relies on word profiles derived from articles citing a cluster of co-cited articles. ... Word-profile approaches have drawbacks. First, word profiles may not converge to a focused message. Analysts and users will make a substantial amount of sense-making efforts to synthesize a diverse range of word profiles. Second, cluster labels based on aggregating word profiles tend to be too broad to be useful. In practice, many users would be interested in not only the most commonly used terms but also terms that can lead to profound changes. Terms associated with an emerging trend could be overshadowed by a broader and more persistent theme.
In CiteSpace II, a current research front is identified based on such burst terms extracted from titles, abstracts, descriptors, and identifiers of bibliographic records. These terms are subsequently used as labels of clusters in heterogeneous networks of terms and articles.
CiteSpace II makes it easier for users to identify pivotal points. In addition to inspecting salient visual attributes, the user easily can see nodes with high betweenness centrality (Freeman, 1979).
The procedure of using CiteSpace II is described in the following steps, 
(1) Identify a knowledge domain using the broadest possible term.
(2) Data collection
(3) Extract research front terms: CiteSpace II first collects n-grams, or terms, from titles, abstracts, descriptors, and identifiers of citing articles in a dataset. The present study used single words or phrases of up to four words. ... Research-front terms are determined by the sharp growth rate of their frequencies.
(4) Time slicing
(5) Threshold selection
(6) Pruning and merging: Pathfinder network scaling is the default option in CiteSpace II for network pruning (Chen, 2004; Schvaneveldt, 1990).
(7) Layout
(8) Visual inspection
(9) Verify pivotal points
We demonstrate the new features of CiteSpace with case studies of two research fields: mass-extinction research (1981–2003) and terrorism research (1990–2003).
Mass-extinction research (1981–2003).
The input data for CiteSpace II were retrieved from citation index databases via the Web of Science based on a topic search for articles published between 1981 and 2003 on mass extinction. The scope of the search included four topic fields in each bibliographic record: title, abstract, descriptors, and identifiers. The search was limited to articles in English only.
The resultant dataset contains a total of 771 records.
A total of 333 research-front terms were detected from the four topic fields of these records.
Terrorism research (1990–2003).
The terrorism research (1990–2003) dataset consists of 1,776 records resulted from a topic search on terrorism in the Web of Science.
A total of 1,108 research-front terms were found.
The fully integrated representation of research fronts and intellectual bases in the same network visualization has three practical advantages.
First, using surged topical terms rather than the most frequently occurring title words is particularly suitable for detecting emerging trends and abrupt changes. In visualized networks, research-front terms are explicitly linked to intellectual-base articles. This design presents a compact representation of the duality between a research front and its intellectual base.
Second, research-front terms naturally lend themselves to be used as labels of specialties.
Third, it overcomes a common drawback of word-profile-based labeling approaches. Aggregated word profiles may not converge to an intrinsic focus. Terms selected based on sudden increased popularity measures are particularly suitable to characterize a current research front.
The Pathfinder algorithm extracts the most salient patterns from a network, but it does not scale well. CiteSpace II implements a concurrent version of the algorithm. The concurrent Pathfinder algorithm has substantially optimized the network scaling module, although it still took 6,000 seconds to process 14 networks and merge them into a 1,704-node network.
In conclusion, the new features introduced to CiteSpaceII for detecting and visualizing emerging trends and abrupt changes in a field of research have produced promising and encouraging results. The major findings are that
• the surge of interest is an informative indicator for a new research front;
• using heterogeneous networks of terms and articles provides a comprehensive representation of the dynamics of a specialty;
• research-front terms are informative cluster labels;
• citation tree-ring visualizations are visually appealing and semantically interpretable;
• betweenness centrality metrics identify semantically valid pivotal points.

2013年12月7日 星期六

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

本研究提出一個TTR-LDA-社群模型,這個模型以推論機制(inference mechanism)結合LDA(Latent Dirichlet Allocation)模型和Girvan-Newman社群偵測(community detection)演算法提供在網路資料上偵測社群並對這些社群進行主題探勘(topic mining)的功能,並且進而了解在社群上的主題隨時間推移的變化,處理的架構如下圖所示

本研究利用Delicious社會標籤系統(social tagging system)上從2005到2008年的資料進行研究。在社群偵測部分,首先建立網絡:根據使用者標籤的資源數量,選取前50000位標籤資源最多的使用者;然後對他們標籤的網頁進行統計,從其中選取10000個被最多使用者標籤的網頁。接著在上述的10000個網頁中,如果有這50000位使用者之間有任何兩位曾經標籤過相同的網頁,便在這兩位使用者之間產生一個連結。以50000個使用者為節點,同時以他們之間的連結為連結線,便可以建立一個共同書籤網絡(co-bookmark network)。並且為了研究網絡上社群結構的變化,並將整個期間的資料分為三個時段:分別為2005-2006、2007與2008年,相關的統計數據如下表:

本研究利用Girvan-Newman演算法找出標籤者(Tagger)社群,使得標籤者與同一社群內的其他標籤者比社群外的標籤者有較強的關係。這個演算法重複移去當時網絡上中介性(betweenness)最大的連結線,產生各種可能的網路劃分(network partition),測量每一種劃分下的群組性(modularity),也就是實際上社群內的成員彼此間的連結線數量與相同連結度的情況但隨機產生連結線的數量的差,群組性最大的劃分便是輸出結果。

另一方面,本研究利用TTR-LDA模型找出每個標籤者的主題分布以及主題內具有代表性的標籤,TTR-LDA模型修改自ACT(author-conference-topic)模式[11][12],是一個由標籤者做為第一層、標籤與資源為第三層、主題則為第二層,所構成的三層貝氏模型(three-layer Bayesian model)。

整合Girvan-Newman演算法和TTR-LDA模型的方法是以社群為單位,將社群內所有標籤者的主題分布進行平均做為該社群的主題分布,根據主題分布,選出機率值較大的主題做為社群的代表。比較不同時段社群共同的代表標籤衡量它們的相似性。

社群偵測的結果發現前五個最大的社群在四年裡占了絕大多數的比率,而且這個比率逐年增加。主題探勘的部分則測量TTR-LDA模型在不同主題數量上的複雜度(perplexity),發現150個主題時有最低的複雜度。因此,以下的研究便針對150個主題在前五個最大社群上的分布進行探討,計算它們的傳導性(conductance)與模組性。就社群模組性而言,最近一個時段(2008年)的結果比前三個時段還要高,其原因可能是因為經過一段時間後,社群的結構逐漸成熟,因此後期比前期更能產生較佳的社群。另外,將最後一個時段再細分為四個較小的時段則發現,較小時段的社群模組性比整年的結果來得高,本研究認為造成這種現象的原因可能是由於在不同的時段,大部分標籤者的書籤行為集中在不同的領域;當那些時段合併起來的時候,會展現標籤者在多個領域的興趣,使得社群內的叢集(clustering)特性較弱。

結果並可以發現前二十個主題都出現在不同時段的前五個最大社群裡,每個社群至少包括一個前十名的主題。此外,比較LDA、TTR-LDA和TTR-LDA-社群等三種模型在資源和標籤上的預測力,在回收率(recall rate)、精確率(precision)和F1指標上以TTR-LDA-社群為最佳。

In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism.

The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008.

Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time.

From a research perspective, these real-world networks display unique properties from the classical random graph model [3] in that most real word networks exhibit three common properties: the small-world property, power-law degree distribution and a high clustering coefficient or transitivity (indicating community structure) [7][8][9].

Thus, an important task in network analysis is to detect communities and explore their features, which can improve community-supporting services at the community-level in the context of a social tagging system.

Many studies in various disciplines have been devoted to community detection; however, few of them have systematically and quantitatively studied the profiles of those detected communities.

In this paper, we propose a TTR-LDA-Community model, which is an inferential combination of an extended LDA model and a betweenness-based community detection algorithm. It provides rich, systematic, and quantitative information about the profiles of detected communities.

In the context of social tagging systems, where multiple users are annotating resources, the resulting topics reflect a shared view of the document; and the tags of the topics reflect a common vocabulary.

Girvan and Newman extended the betweenness measure to edges and designed a clustering algorithm which gradually removes the edges with the highest betweenness value [4]. This algorithm has been improved through modularity; and the complexity is reduced from O(m2n) to O(mdlogn) where d is the depth of the dendrogram of the community structure [2].

Many studies provide various models and algorithms for topic mining and community detection; yet, few of them have integrated those models and algorithms, performed topic mining for detected communities, and analyzed how those identified topics change among communities over time.

The activity of social tagging consists of three major components: tag, tagger and resource. The experimental dataset contains all the triples of these three components and the time and date of their creation on Delicious from 2005 to 2008.

In data processing, all taggers were ranked by the number of resources they have bookmarked and the top 50,000 taggers were selected as the sample of taggers.

These taggers bookmarked a total of 354,522 web pages, which were sorted by the number of taggers who bookmarked them. The top 10,000 resources were selected as the sample of web pages, associated with which a dominant majority of tagging activities occurred.

Thus a co-bookmark network was built in which a connection between two users (within the sample of 50,000 taggers) is created if they bookmarked the same resources (within the sample of 10,000 web pages).

In addition, in order to observe the evolution of structure and motif of communities, the time span (2005-2008) was divided into three slices.


The model is illustrated in Figure 1. TTR-LDA is developed based on ACT model [11][12]. It is a three-layer Bayesian model with taggers tap in each post p as the first layer, tags t, and resource r as third layer and all the topics denoted as latent variable z as the middle layer.

The inference mechanism is used to infer the topic distribution over detected communities.

Each community includes a set of taggers, who have a stronger relationship with other taggers within the community than the taggers outside.

Based on the taggers’ information model, the probability distribution of each tagger over a set of topics is obtained by using the TTR-LDA model while the community structure of taggers is revealed by the community detection algorithm. The two sets of results are further integrated through an inference mechanism.




Results show that the number of users of the top five communities occupies a major proportion in the four years (2005-2008) and the proportion is increasing over time.

Perplexity is used to identify the number of topics [10], which arrives at the lowest point when the number of topics is 150. The interest model of each tagger in the top five largest communities is then built based on their topic distributions.

By using users’ interest models and the inference mechanism, a topic distribution of the largest community can be created (Figure 3). We can find that the topic distributions in a community are diverse because users’ relationships in that community are mainly based on their co-bookmark activities not the similarity of their interest model.

In order to observe the dynamic features of communities, we design an experiment as follows:
1) denote the five largest communities from each time slice in 2008 as community_i_t where t means the tth time slice in 2008 and i means the ith largest community in tth time slice;
2) compute the topic distribution for the five communities, which is stored as model_t_i_Topic(j), the
probability of jth topic in ith largest community in the tth time slice;
3) obtain the probability distribution of tags that are collected from all the posts generated during the specific time slice; the probability of one tag occurring in a topic shows the level of representativeness of the tag for that topic;
4) sort all the tags according to their probability value in each topic and select the 20 top ranked tags to represent the content of the topics; select the top 5 ranked topics to represent the theme of each community;
5) analyze the similarity between different communities from different time slice through computing how many tags are shared by the two different communities. More specifically, we compare current time slice with its previous time slice, for example, we compare community_i_t with community_j_t-1 (j=1, 2…5).

The size of communities along evolutionary lines fluctuates over time. For example, the size of the community about social networks in the 3rd time slice (community_1_3) is much larger (4,377) than that (521) in the 4th time slice (community_5_4).

Conductance (from multi-criterion scores) and modularity (from single criterion scores) are used to evaluate the quality of communities detected by the TTR-LDA-Community model [6].

The smaller the value of conductance is, the higher the granularity of a community is. Network community profile (NCP) is used to compute and display the value of conductance for communities [5].

Whiskers networks and rewired networks are adopted as two comparative aspects. Whiskers is defined as the maximal sub graphs that can be detached from the rest of the network by removing a single edge; and a rewired network is a random network that has the same nodes and the same degree distribution as the original network [5].

The conductance of communities of the rewired original network (blue line in the left figure), rewired random network (red dashed line in the left figure), the original whiskers network (blue line in the right figure), and the random whiskers network (red dashed line in the right figure) are calculated and shown in Figure 4.

In Figure 4, compared with the rewired network (left) and the rewired whiskers (right), 1) the original network displays a higher granularity of communities (a lower conductance value); 2) the value of conductance as the function of the size of communities in the original network and the original whiskers present a “V” shape, showing properties of a true large social networks [5]; 3) the original whiskers has the best community granularity (the lowest conductance) between size 10-100; and 4) the best community granularity of rewired original network is around 1000.

The modularity of communities in the four time slices of 2008 is better than that in 2005-2007. This is probably due to the fact that community structure grows mature gradually over time, creating better communities in later years than in earlier years.

Meanwhile, modularity of communities in the short-term (four sub periods in 2008) is larger than the long-term (2008). It can be explained that in different time periods, most taggers’ bookmarking activities are focused on different domains, so in a certain short-term time period, communities may be quite different from each other. However, when those time periods are merged together, the taggers show different interests in many domains; so the clustering feature within the communities becomes weaker.

Results show that the most popular topics are about bandslash fiction, fan fiction, and supernatural fiction (the top 3 popular topics). Communities with similar theme are ranked 3rd, 4th, and 5th in size; and the web resources with similar topics are ranked 500-600 of the top 1000 ranked resources in number of taggers associated with them.

The top 20 ranked topics in 1000 most popular resources can be found in 5 largest communities in different time periods. For each community, there exists at least one topic that is ranked top 10 in 1000 most popular resources (Table 3).

Topic distributions for each community are obtained respectively from LDA, TTR-LDA model, and TTR-LDACommunity model based on co-bookmark network in a given period (Oct. 2008–Dec. 2008). One resource and five tags are recommended for each post according to the results of three models separately.

The TTR-LDA and TTR-LDA-Community model show significant improvement for recommendation of tags and resources for post in terms of precision, recall and F1-Measure. TTR-LDA and TTR-LDA-Community have slightly improved performance for “tags for post”, while TTR-LDA-Community outperforms TTR-LDA on “resource for post”.


2013年12月3日 星期二

Rzeszutek, R., Androutsos, D., & Kyan, M. (2010). Self-organizing maps for topic trend discovery. Signal Processing Letters, IEEE, 17(6), 607-610.

Rzeszutek, R., Androutsos, D., & Kyan, M. (2010). Self-organizing maps for topic trend discovery. Signal Processing Letters, IEEE, 17(6), 607-610.

就多筆文件資料以及多個詞語的語料庫而言,可以定義一個詞語-文件的共現矩陣(co-occurance matrix), C,紀錄每個詞語在每筆文件中的出現次數。LDA將矩陣C分解成兩個矩陣Φ和Θ。Φ表示主題-詞語的可能性,其中Φ的第k個向量ϕk中的每個元素即為每個詞語在第k個主題上的出現分布。Θ則是文件-主題可能性,它的第d個向量θd上的元素代表文件d包含各主題的可能性。進一步來說,Θ定義了一個文件空間(document space),這個空間上的每一個維度描述相對應主題在文件上的重要性,當文件彼此在意義上相似的話,他們在文件空間上也會相當接近。如果將研究的文件資料依照發表時間區分成若干的時段,每一個時段tw的主題分布情形θ(tw)可以用這個時段內所有文件的主題分布情形的平均值代表,如下面的式子



在這裡,ND(tw) 是在時段tw內發表的文件數量。過去 [5]便曾經利用LDA對於科學論文進行主題模型的研究,從產生的圖形表現出全球暖化研究在十年間逐漸增加的趨勢。然而有時主題的數量多達100個,若是同時將所有的主題在時間上的變化顯示在圖形上,在分析上便很難觀察每一個主題變化的模式(patterns)。
本研究建議將 自組織映射圖(self-organizing map)的視覺化功能結合LDA模型以追蹤主題在時間上的變化情形,藉由觀察自組織圖上反應(response)模式的變化,可以詳細地了解資料集中內容的改變情形。Kohonen的自組織映射圖 [8]能夠將高維度的資料映射到較低維度的格狀(lattice)圖形上,因此本研究嘗試運用這項技術來解決主題數量龐大時的視覺分析問題。自組織映射圖包括訓練(training)與分群(classification)兩個階段。訓練時重複地隨機選取文件資料,比較每一個節點與文件的相似程度,選取最相似的節點,使選取的節點與鄰近地區的節點都往文件的主題方向調整,訓練階段完成後便能使自組織圖代表全部的文件空間,而每一個節點則能代表主題相似的節點。有別以傳統使用所有資料來訓練自組織圖,本研究僅從文件資料集中隨機抽取500筆文件進行訓練,以減少計算的負荷,並且更大的不同是以KL差異(Kullback-Liebler divergence, KL divergence)取代常見的Euclidean距離(Euclidean distance)做為比較文件主題分布與節點主題分布的相似程度。兩個文件在文件空間上的向量分別是A與B時,它們的KL差異定義為
然而KL差異不符合三角形不等式(triangle inequality)與對稱律,也就是KL(A, B) <> KL(B, A)。為了讓文件間的相似性測量具有對稱性,本研究建議使用下面的式子

分群時將資料集上文件依照發表時間映射到對應時段的自組織圖上,因此每一時段會產生一個自組織圖的分群結果,對圖上的每一個節點nj定義一個分群密度 D(nj)

此處Nj是該節點nj上的文件數量,NmaxNmin分別是該時段分群結果的自組織圖上所有節點的文件最大與最小數量。最後根據這些結果產生二維圖形。針對個別時段的自組織圖分群圖形可以發現該時段內的重要主題,連續觀察所有時段的圖形則可以看出各種主體的變化情形。

本研究以2009年五月到八月ESPN.com與TSN.ca的25754筆體育新聞與評論的RSS feed為分析資料。

In order to track changes in topics over time, simple time-series techniques have been applied to a corpus analyzed using LDA [5]. For instance, in [5], the authors show how scientific papers on global warming gradually increase in popularity over a ten year period. Unfortunately it is not uncommon to have 100+ topics (i.e., dimensions) in a descriptor which makes this analysis difficult for any more than two topics.

Therefore, we propose to use a method similar to the WEBSOM [6] and ProbMap [7] algorithms to perform the trend analysis. These algorithms use Kohonen’s Self Organizing Maps [8] to nonlinearly project a high dimensional feature space onto a low-dimensional output space. Our method merges the idea behind WEBSOM and ProbMap with the work done in [5] to show how a document corpus can change with time.

Therefore, given a corpus, it is possible to define a word-document co-occurance matrix, C, that relates the number of times a given word occurs for a given document. It is desirable to find a way to decompose C because a corpus may contain millions of documents and thousands of words.

Latent Dirichlet Allocation probabilistically decomposes C such that


where Φ and Θ are matrices that express how words, topics and documents are related. Φ contains the topic-word likelihoods or, put another way, how likely any given word is to appear in topic k. Θ is a collection of document-topic likelihoods which relates how likely a document is to contain topic k.

Therefore, for each topic k, there is a vector, ϕk, that contains the word distribution for that topic. Similarly, there exists a vector θd for each document d that contains the topic distribution for that document.

The topic-distribution matrix, Θ, acts as a natural descriptor for all of the documents in the corpus. For any document, d, its associated vector θd  describes its location in a document space.

LDA ensures that documents that are semantically similar (i.e., share many words that are in the same topics) will be close to one another in the document space.

A common choice of divergence in this sort of situation is the Kullback–Liebler (KL) divergence and it is defined for two-dimensional PMFs, A and B as

The KL divergence is not a distance measure (i.e., metric) so it does not satisfy the triangle inequality and KL(A, B) <> KL(B, A). For convenience, it is useful to use a symmetric KL divergence so that the order of the arguments is not important. The symmetric KL measure is defined as


For this paper, a corpus was constructed of 25 754 documents that were collected over a three-month period starting in late May 2009 and ending in late August 2009. The documents were article summaries from RSS feeds from sports news and opinion websites such as ESPN.com and TSN.ca.

The values of the hyperparameters were taken from [5] and NT = 40 since we found that this number of topics provided the best tradeoff between descriptor length and the ability to effectively describe the corpus.

The moving average approach simply produces a document descriptor that is the average descriptor for any particular time window. For each window, tw, we obtain a descriptor, θ(tw), such that



where ND(tw) is the number of documents inside of the time window at time tw. θ(tw) then represents the central tendency of the documents inside of that time window.

Unfortunately, it is very difficult to extend this sort of analysis to more than just two or three topics. Consider Fig. 3. All of the topics are shown on the same plot and it is extremely difficult to visually see any underlying patterns. More importantly, this type of trending only shows the most dominant document types at any point in time.

The Self-Organizing Map (SOM), as proposed by Kohonen [8], maps a high-dimensional feature space onto a lower dimensional representation (usually one or two dimensions).

This allows the map to perform a nonlinear dimensionality reduction on a dataset. However, it can also be used as a classifier, which is what we do in this paper. By examining how many data points each node classifies, it is possible to map complex structures in the feature space onto the lattice.

The first stage trains the SOM on a small, random subset of the original input data. For the dataset used in this paper, that subset is 500 randomly selected documents. This is done since we do not need the SOM to accurately model the entire dataset, just loosely resemble it. This has the added benefit of reducing the computational burden when training the SOM, especially for very large datasets. After training, the SOM will now resemble Θ such that each node is actually the representation of a cluster of similar documents (Fig. 4).

As discussed in Section II-B, KL-divergence is used for the “distance” measure since it well suited to describing the dissimilarity between the document probability vectors.

The second stage filters, or processes, the dataset through the SOM using the sliding window method described in Section III-B. We use the trained SOM as a classifier to determine how many documents in the window are classified by each node. Each node has an associated classification count, Nj, or the number of documents that are associated with node j. A document is associated with a node if nj is the closest node to that document vector.

We define a classification density, D(nj), such that



where Nmax and Nmin are the maximum and minimum count values in the window. This ensures that the map is normalized to be the range of [0,1] so that different windows can be compared.

As the map responses change over time, it defines a 3-D volume (Fig. 7). This volume describes how the map responds over time, as opposed to just observing the response of the SOM at any particular time. As before, the clustering properties of the SOM makes it possible to determine how the document distribution itself changes.

2013年11月20日 星期三

Sugimoto, C. R., Li, D., Russell, T. G., Finlay, S. C., & Ding, Y. (2011). The shifting sands of disciplinary development: analyzing North American Library and Information Science dissertations using latent Dirichlet allocation. Journal of the American Society for Information Science and Technology, 62(1), 185-204.

Sugimoto, C. R., Li, D., Russell, T. G., Finlay, S. C., & Ding, Y. (2011). The shifting sands of disciplinary development: analyzing North American Library and Information Science dissertations using latent Dirichlet allocation. Journal of the American Society for Information Science and Technology, 62(1), 185-204.

本研究利用北美在1930到2009年間完成的3121筆博士論文,探討圖書資訊學(library and information science, LIS)主要研究主題的變化情形。過去在研究LIS領域的主題時,內容分析(content analysis)和共被引分析(cocitation analysis)是主要的研究方法,針對一個時期內的期刊論文進行分析,發現該時期主要的研究主題。重要的內容分析研究例如 Enger, Quirk, & Stewart (1989)、 Järvelin & Vakkari (1990, 1993)、Kumpulainen (1991)、 Hider & Pimm (2005)和  Fidel (2008);以期刊進行書目計量分析的研究有 Åström (2007, 2010)的兩篇論文;以論文作者進行書目計量分析的研究則包括Pettigrew & Nicholls (1994)、White & McCain (1998)、 Bates (1998)、 Budd (2000)、White (2001)、Levitt & Thelwall (2009a, 2009b)和Åström (2010)。目前大多數的研究有下列的限制:1)雖然大多數的研究都以期刊論文為研究資源,但已有許多研究(Bazerman, 1988; Hyland, 2000)指出不同文類(genres)的書寫與引用模式皆有不同,只有針對一種文類進行分析便很有可能只會產生單一觀點;2)除了以期刊論文為主要的研究資料外,大多數的研究並且以少數高被引的作者或隨機選取的論文做為代表性的資源,但由於分析的資料數量不大,容易受到少數的影響,使產生的結果可能無法代表全體的情形;3)目前的研究大多是針對一個時期的同時性(synchronic)研究,缺乏以趨勢變化分析為主的歷時性(diachronic)研究,少數的歷時性研究有Smeaton, Keogh, Gurrin, McDonald, & Sødring (2003)針對SIGIR研討會論文探討資訊檢索領域在25年間的變化,Sugimoto & McCain (2010)同樣探討資訊檢索領域的主題發展情形,Harter & Hooten (1992) 將1972-1990年分為三個時期研究the Journal of the American Society for Information Science & Technology articles上的論文與作者在書目計量上的變化,Järvelin and Vakkari (1993)則是分別對三個時期的期刊論文進行內容分析, Åström’s (2007) 將 LIS領域相關的期刊論文分為三個時期進行共被引分析。除了這些限制以外,上述研究的另一個問題是這種方法間接透過論文與作者進行分析,並非直接針對主題進行研究。先前Braam, Moed,&van Raan (1991a, 1991b)的研究嘗試利用詞語的共同出現做為分析的資訊來探討領域內的重要研究主題,然而Leydesdorff (1997)認為在不同的文件累積範圍內,單一詞語的出現有相當不同的意義,因此這樣的分析有待商榷。

本研究的研究方法是以LDA(latent Dirichlet allocation) (Blei, Ng, & Jordan, 2003)來確認各時期的隱藏主題,並且確認各個主體的代表論文。LDA是一種文件生成模型,目前已被廣泛地應用於主題的確認與分析,例如Rosen-Zvi, Grifftihs, Steyvers,&Smyth (2004)曾運用LDA探討作者和主題之間的關係;Tang, Jin, & Zhang (2008)延伸這樣的概念到學術網路(academic networks)上;McCallum,Wang, & Corrada-Emmanuel (2007)和Li et al., (2010)則分別運用LDA於社交網絡(social networks)和社會標籤社群(social tagging communities)。Blei & Lafferty (2007)則運用LDA了解各主題彼此間的相關性(correlations),Pruteanu-Malinici, Ren, Paisley,Wang, & Carin(2010)和Rzeszutek, Androutsos, & Kyan (2010)的研究關注在主題在時間上的變化。為了了解LDA進行主題確認的可行性,Griffiths & Steyvers (2004)和Zheng, McLean, & Lu (2006)將LDA產生的主題與現有的論文資料分類進行比較,也證明了這種方法的確可行。 LDA將文件表示成由隱藏的主題隨機混合而成,也就是每個文件包含多個詞語,並且每一個主題由詞語依不同比例分布。文件上的所有詞語是從這個文件對應的主題組合中依據它們的比例重複地隨機抽取而得,配合主題對應的詞語比例選出。本研究使用的LDA方法是由Rosen-Zvi, Grifftihs, Steyvers,&Smyth (2004)所擴充的作者-主題模型(author-topic models),在這個模型中不僅允許一個文件可以包含多個主題,它描述了一個文件可能具有多位作者,並且每位作者也可以針對多個主題書寫的情形。下圖是作者-主題模型的Bayesian網路模型示意圖:

1. 對某一位作者x來說,Θ 是選擇某一個主題z的機率分布。
2. 對某一個主題z來說,φ 是選擇某一個詞語w的機率分布。
3. ad 是文件的多位作者,x是從這些位作者當中隨機選取的一位。

在估算上述LDA模型裡的給定某一位作者後選擇某一個主題的機率(Θ)以及給定某一個主題後選擇某一個詞語的機率(φ)等兩個未知參數時可以使用Gibbs取樣演算法(Gibbs sampling algorithm),這個演算法使用連續Markov鏈取樣(successive Markov chain sampling),根據所有其他變數,重複抽取一對的作者(x)與主題(z),以實際上所得到的數值來估算

此處,nw[m][j]是第m個詞語被指定給第j個主題的次數,nkwsum[j]是所有詞語被指定給第j個主題的次數總和,V是詞彙的大小,也就是語料內共有多少種詞語種類的數量,na[x][j]是第j個主題被指定給作者x的次數,naksum[x]是所有主題被指定給作者x的次數總和,T是主題的數量。
根據上面的過程,φ與Θ的估算方式如下


在實際研究上,可以根據複雜度(perplexity)測量模型的成效來選擇主題的數量,愈小的複雜度表示模型的成效愈好。

本研究將1930-2009年區分為五個時期,每一個時期設定為50個主題,選擇該時期機率較大的五個主題視為是該時期的主要研究主題,並且對每個主題選出機率較大的詞語來了解該主題。研究發現如下圖

從開始時期(1930-1960)到現在(2000-2009)LIS的主題有本質上的改變,然而也有圖書館史(library history)、引用分析(citation analysis)與資訊檢索(information retrieval)等出現在多個時期內,這些可視為LIS的核心主題。

This work identifies changes in dominant topics in library and information science (LIS) over time, by analyzing the 3,121 doctoral dissertations completed between 1930 and 2009 at North American Library and Information Science programs.

The authors utilize latent Dirichlet allocation (LDA) to identify latent topics diachronically and to identify representative dissertations of those topics.

The findings indicate that the main topics in LIS have changed substantially from those in the initial period (1930–1969) to the present (2000–2009). However, some themes occurred in multiple periods, representing core areas of the field: library history occurred in the first two periods; citation analysis in the second and third periods; and information-seeking behavior in the fourth and last period.

Many evaluations of library and information science (LIS) have been conducted, primarily using the methods of content analysis and cocitation analysis on journal articles (e.g., Järvelin & Vakkari, 1993; White & McCain, 1998).

Although these studies constitute one lens on the field, there are some major limitations to the current literature in the area.
First, the focus on a single communicative genre (the journal article) provides a monocular view of the field. Research has shown that the writing and citing patterns of authors vary significantly by genre (Bazerman, 1988; Hyland, 2000). A different topic spectrum may be found by examining topics across multiple genres.
Second, the focus has been on either a group of highly cited authors or a sample of journal articles. Previous analyses have been manually intensive, necessitating small sample sizes. This has the potential to skew the results in two ways: (a) highly cited works are not necessarily representative of the works produced, and (b) a few articles/authors can heavily influence the results.
Lastly, the analyses have been largely synchronic, rather than diachronic. Therefore, trend data rely on replication studies, which are not prevalent in the literature.

Many quantitative analyses have been conducted to analyze the domain of LIS: content analyses of journal articles (e.g., Enger, Quirk, & Stewart, 1989; Fidel, 2008; Hider & Pimm, 2005; Järvelin & Vakkari, 1990, 1993; Kumpulainen, 1991), bibliometric analyses of journal articles (e.g., Åström, 2007, 2010), and bibliometric analyses of authors (e.g., Åström, 2010; Bates, 1998; Budd, 2000; Pettigrew & Nicholls, 1994; Levitt & Thelwall, 2009a, 2009b; White, 2001; White & McCain, 1998) to provide large-scale descriptions of the field.

Some analyses have focused on particular journals (e.g., Harter & Hooten, 1999; Lipetz, 1999; Liu, 2002; Park, 2010), conference proceedings (e.g., Smeaton, Keogh, Gurrin, McDonald, & Sødring, 2003), subject areas (e.g., Sugimoto & McCain, 2010), or countries (e.g., Cano, 1999; Uzun, 2002).

Scholars have also performed bibliometric analyses to examine the relationship between LIS and other disciplines (e.g., Borgman & Rice, 1992; Ellis, Allen, &Wilson, 1999; Meyer & Spencer, 1996; Odell & Gabbard, 2008; Sugimoto, Pratt, & Hauser, 2008).

The majority of quantitative analyses of LIS share four things: (a) the journal article is the focal communicative genre, (b) they are synchronic, rather than diachronic, (c) they focus on relationships between journals and/or journal authors (rather than topic analysis), and (d) those focusing on topic analysis use methods of co-occurrence or content analysis.

Some notable exceptions (on at least one of the four points) are the works by Smeaton et al. (2003) and Sugimoto and McCain (2010), which looked at changes in topics in information retrieval over time; Harter and Hooten’s (1992) bibliometric study of the Journal of the American Society for Information Science & Technology articles for three time slices; Åström’s (2007) cocitation analysis of LIS journal articles for three periods; and Järvelin and Vakkari’s (1993) content analysis of journal articles for three periods.

Some work has been done to examine the value of various methods of topic analysis, comparing the results found through cocitation with co-word analysis (e.g., Braam, Moed,&van Raan, 1991a, 1991b) and the difference between using titles, author-supplied, or indexer-supplied keywords (Whittaker, Courtial, & Law, 1989). ... Scholars have also criticized the use of co-occurrence analysis of terms, noting the large variance in meanings of individual terms based on the level of textual aggregation under investigation (Leydesdorff, 1997).

However, the manual intensity of content analysis becomes difficult, as the number of dissertations in the discipline has gone from a few hundred to a few thousand. Although content analysis can provide high granularity for individual works, it becomes difficult to assess the entire body of work without automatic techniques.

Latent Dirichlet allocation was proposed by Blei, Ng, and Jordan (2003) as a generative probabilistic model useful for discovering underlying topics in collections of data.

Expansions of LDA have also been used to understand correlations between topics (Blei & Lafferty, 2007), authors (Rosen-Zvi, Grifftihs, Steyvers,&Smyth, 2004), academic networks (Tang, Jin, & Zhang, 2008), social networks (McCallum,Wang, & Corrada-Emmanuel, 2007), social tagging communities (Li et al., 2010), and changes in topic over time (Pruteanu-Malinici, Ren, Paisley,Wang, & Carin, 2010; Rzeszutek, Androutsos, & Kyan, 2010).

Two exceptions to this are Griffiths and Steyvers’ (2004) analysis of abstracts from the Proceedings of the National Academy of Science (PNAS) from 1991–2001 and Zheng, McLean, and Lu’s (2006) analysis of the bioinformatics literature (from MEDLINE abstracts). These studies found that LDA performed well, in that the latent structure mimicked some characteristics of explicit structure (such as categorization schemes). Moreover, the studies displayed the ability of LDA to analyze the rich underlying structures of the domain— depicting emerging and sustained trends in a given discourse.

In LDA, a topic is characterized by a distribution over words and then documents are represented as random mixtures over latent topics (p. 996). As a three-level hierarchical Bayesian model, each topic node is sampled repeatedly. The result is that words may be repeated within topics and documents may be associated with more than one topic.

This model was extended to what is called the author–topic model (Rosen-Zvi et al., 2004), which is used in the present analysis. In this model, not only is each document a mixture of probabilistic topics, but each author is also seen as a mixture of probabilistic topics. In the same way that a topic can be generated from multiple topics, this extended model recognizes that an author can be “about” multiple topics (within a single document or across documents).

The author–topic model allows us to not only examine which topics were most salient across the various periods, but also which authors are most associated with these topics.



The figure can be explained as follows:
1. Θ is the probability of a topic given an author x; α is a hyperparameter for Θ.
2. φ is the probability of a word w given a topic z; β is a hyperparameter for φ
3. ad provides for the fact that multiple authors can write a single document; x is a randomly selected author from ad. (Note that there are only single authors in this selection; however, it is still necessary to identify author x.)
4. Given author x, we identify the topic z most likely to be associated with the given author.
5. Given topic z, we identify the words w most likely to be associated with the given topic.

Given the Bayesian network created from Figure 1, the joint probability of the author–topic pair was estimated using Gibbs sampling algorithm (Casella, 2001). This algorithm allows for an estimation of the unknown parameters of the model, namely, the probability of a topic given an author and the probability of a word given a topic. Gibbs sampling uses a successive Markov chain sampling to repeatedly draw x and z (as a pair), conditioned on all other variables. The process can be expressed as follows:



where nw[m][j] is the number of times a single word is assigned to a topic; nkwsum[j] is the number of times any word is assigned to a topic; na[x][j] is the number of times a single author is assigned to a topic; naksum[x] is the number of times any topic is assigned to an author.

Once this process has been iterated 2,000 times, the results of these variables are used to calculate Θ and φ.


Perplexity analysis is used to estimate the performance of the model. This analysis is used when we have an unknown probability distribution in the data. A lower perplexity value indicates better performance. As shown in Figure 2, the performance for our data stabilized at 50 topics.

At this point, all data were analyzed according to the author–topic model, divided into the five time slices. For each period, 50 topics were identified. Each topic contained a probability value—that is, the likelihood that the topic identified should be associated with the period. These topics were ranked by probability values and the top five were selected as being most representative of the period.

Similarly, a probability for each word was calculated to represent the association between a word and the given topic. These were ranked by probability values and the top 20 were chosen as most representative of the topic.

Lastly, the authors were assigned probability values for each topic and these too were ranked. The top five were chosen as highly representative authors for the given topic.

The results of the topic analyses are summarized in Figure 3, with the topics ranked from highest (top) to lowest for each period.



Some topics occur across multiple periods, including library history, librarianship, information use, citation analysis, classification, information retrieval (abbreviated as IR in Figure 3), and information-seeking behavior. These are the core topics in LIS dissertations from 1930 to 2009.

A limitation inherent with LDA analysis is in the manual interpretation and labeling of “topics.” Although some topics were fairly straightforward to label (e.g., Topic 5c, the top three loading words of which were (a) information, (b) seeking, and (c) behavior), others proved more difficult to ascertain the content or methodological relationship that connected the words and dissertations.

The top three topics by period identified in Järvelin and Vakkari’s (1993) content analysis of LIS journal articles is displayed in Table 7. These topics share some similarity with the present findings: classification is certainly a top theme in the first period, but there is an overwhelming emphasis on history in this period that is not represented by Järvelin and Vakkari’s results. ... The discrepancies between these findings lead to questions relative to genre and the possible impact of genre upon the topics. Further analysis should explore whether certain topics appear first in particular communicative genres and whether genres consistently emphasize different topics within a single field.



In terms of the highest loading specialties, this work confirmed an interest in information retrieval, bibliometrics, and information use consistent with the findings of White and McCain. However, the current analysis places a stronger emphasis on library evaluation, management, and education than was the case with White and McCain’s study. Although White and McCain did not undertake a detailed analysis on the changes in topics over time, they provided evidence for a shift toward the cognitive and user side of information retrieval.

Åström (2007) similarly evaluated LIS using cocitation analysis. He identified the main clusters by period as shown in Table 8.



For example, numerous studies summarized above noted the lack of theoretical work published in the LIS literature. It may be the case the journal articles are more likely to cover experimental work, whereas dissertations tend toward the theoretical.

However, the findings suggest that the bulk of dissertations completed at these schools do not have explicit connection to library practice.

Houser (1982) states that a “discipline is formed to solve a range of problems about some natural or social phenomena. . . [t]hese problems have a genealogy, that is, a continuity which forms the domain of the discipline” (p. 97).

Turner’s (2000) definition of a discipline focused on the identity of a shared name for a specialization and the exchange of scholars within that discipline (thereby propagating the discipline).

In examining dissertations, we were able to identify some dominant themes in LIS: These can be broadly defined as information-seeking, use, access, organization, and retrieval; and the education and training of the professionals providing these services.