顯示具有 topic models 標籤的文章。 顯示所有文章
顯示具有 topic models 標籤的文章。 顯示所有文章

2017年12月7日 星期四

Nichols, L. G. (2014). A topic model approach to measuring interdisciplinarity at the National Science Foundation. Scientometrics, 100(3), 741-754.

Nichols, L. G. (2014). A topic model approach to measuring interdisciplinarity at the National Science Foundation. Scientometrics100(3), 741-754.

跨學科研究(IDR)是指由整合多個學科的理論、技術、資料與工具來解決單一學科無法解決的問題。在IDR的測量時,一般假定所有科學之間是一個連貫的學科結構(a coherent disciplinary structure),並且IDR的表現正是本質上模糊了學科之間的界限(Wagner et al., 2011)識別和測量IDR需要在一個研究計畫中評估多個學科的存在和整合,並且包括評估科學的投入,產出和過程(Wagner et al., 2011)。

測量跨學科性(interdisciplinary)的主要挑戰來自確定和界定構成IDR的不同學科。識別、理解和測量IDR需要對引導出研究方法、理論和結論的知識基礎(intellectual bases)進行解析和特徵化。研究人員已經嘗試了多種方法,但測量跨學科性及其隨著時間的動態仍然是一項艱鉅的工作。質性方法通常利用參與觀察、訪談和調查來描述多學科研究人員團隊中的過程和關係,並評估學科整合的程度(參見Masse et al, 2008; Stokols et al, 2003)。量化方法則通常仰賴於文獻計量學和網路分析技術(參見Porter和Rafols 2009; Rafols和Meyer 2008; Leydesdorff 2007),檢驗論文參考文獻列表中出現學科的引用分析是最常用的方法之一(Wagner et al., 2011)。許多有關測量IDR的文獻計量學文獻都側重於研究科學或出版物的產出(Wagner et al., 2011)。

本研究則是使用美國國家科學基金會(National Science Foundation, NSF)獎勵資料庫(award databse)的給獎建議與獎勵,從範圍更廣泛的人力、投入和過程等方面來測量IDR,描述IDR的互動與整合。

Gerrish和Blei(2010)有關測量學術影響(scholarly impact)的研究,比較了傳統的引文分析和基於語言的主題模型方法。他們發現,雖然這兩種方法在整體學術影響方面有一致的結果,但是基於語言的方法通常能確認在質量上有不同的有影響力的文章。Wang等(2011)結合LDA主題模型與網絡分析技術,開發幫助研究人員評估龐大且迅速增長的生物醫學文獻,以確定化學物質、基因和對藥物發現重要的疾病之間的顯著關聯的工具

本文利用NSF主題模型和NSF的體制結構(institutional structure),探討測量NSF獎勵組合中IDR的新方法。NSF主題模型(Newman et al., 2011)幫助NSF的工作人員與科學界更加了解NSF基金組合的內涵與脈絡,同時也提供文件學科內容(disciplinary content )的新評估方式2000年到2011年間由NSF頒發的獎項約170,000利用這些文件訓練NSF主題模型,共1000個主題,在扣除一些僅包含語言中常用的停字詞所組成的主題之後,共923個,並且依據主題對文件上的關連性,對每個獎項指定一到四個主題,主題的次序代表它們的關連性高低。

本研究利用前述運用主題模型方法產出的獎勵的指定主題,評估SBE(Social, Behavioral, and Economic Sciences)部門管理獎項跨學科

本研究利用NSF所屬的各部門代表學科,MPS (Mathematics and Physical Science) 因為包含多個彼此分離的學科,所以再細分第一步先對923個主題,利用所有170,000個獎項的指定結果以及獎項所屬NSF部門歸類。計算每個部門管理的獎項中每個主題出現的頻率,將主題指定給出現頻率最高的學科。如果有某一個主題高頻率地出現在多個部門或是被分配到非研究或是跨學科的部門,此時則進行個別檢視,根據主題描述將其指定給一個學科或是標示為「非學科特定」。非學科特定的主題例如,t3的假設(Hypothesis)、t60的儀器(Instrumentation)、t738的創業(Entrepreneurship)和t889的研究生(Graduate Students)。根據獎勵和主題在所有部門的分布統計,除了生物學(Biology)和地球科學(Geosciences)以外,其他部門兩者間的分布相當類似。生物學擁有10%的獎勵,但卻有18%的主題被指定給生物學,其原因可能是因為生物學包含多個次學科,而且各自使用相當專殊化的語言來描述他們的科學。反之,地球科學主管NSF23%的獎勵,卻只有6%的主題,其原因可能包括地球科學具有比較狹小的學科範圍、比較仰賴共同語言或者較強的跨學科連結。

本研究針SBE(Social, Behavioral, and Economic Sciences)部門管理的獎項進行跨學科性評估,因此在指定主題對應的主要部門後,再進一步針對SBE在2000到2011年間共有14,225個獎項,通過它們上面出現的主題所屬的學科數量以及主題在獎項上的出現排序,計算它們的跨學科性。如果獎勵包含主題的學科有一個被指定為SBE上其他的學科,則將該獎勵視為「內部跨學科性」(internal interdisciplinarity);若是該獎勵中只要有一個主題屬於其他部門或MPS的學科,則視為「外部跨學科性」(external interdisciplinarity),否則便是無跨學科性。除了上述簡單的三元化數量分析外,本研究也利用Stirling’s (2007)的多樣性指標(diversity index)評估每個獎項的跨學科性

此外,為了比較不同組合的科際整合性,本研究特別挑選6個核心計畫(core programs)進行分析: 社會學(Sociology)、政治學(Political Science)、經濟學(Economics)、地理空間學(Geography and Spatial Sciences, GSS)、決策、風險與管理科學(Decision Risk and Management Science, DRMS)和知覺、行動和認知科學(Perception, Action, Cognition, PAC),分析每個計畫內獎項組合的跨學科與Stirling多樣性指標,並且利用Sci2 Team (2009)的Science of Science Toolkit製作每個計畫的共現網路圖(co-occurrence network diagrams),對於從跨學科的各面向(數量、平衡與差異性)解釋和描述了不同類型的跨學科互動情形

研究結果發現,根據簡單的三元化數量分析,89%的SBE獎項是屬於跨學科研究,外部跨學科性和內部跨學科性分別占55%與34%,如果加上獎項的金額做為加權的話,有93%是跨學科研究,其中外部跨學科性高達74%,而內部跨學科性則是19%。其原因是獲得高額的獎勵大多是外部跨學科研究(約占79%),因為這些研究通常是需要較大成本與跨部門研究團隊的大型計畫



研究結果也顯示每年各類型(內部性、外部性)的跨學科研究數量和學科組合相當穩定,雖然各年之間主題分布有差異。而各年Stirling多樣性指標的平均值則是穩定而小幅成長,其原因可能是由於這些計畫大多為多年性的延續計畫。


六個核心計畫的跨學科研究獎勵比例與它們的平均Stirling多樣性指標有很大的差異,以結果來看,GSS、DRMS和PAC在獎勵比例較Economics和Political Science為大,平均Stirling多樣性指標也有同樣的結果,然而Sociology雖然跨學科研究的獎勵比例較大,然而它的Stirling多樣性指標卻較小。根據Stirling多樣性指標的計算方式,推斷造成Sociology在這項指標上較小的原因可能是在Sociology計畫內雖然許多學科都有跨學科研究的關係,但這些學科大多是SBE內的學科,只有少數SBE外的學科。此外,令一個可能的原因是Sociology內獎勵上使用的語言較一般,許多術語也常出現在其他學科中,所以僅僅只有兩個主題被歸類在這個學科,大部分相關的主題都被指定為非特定的SBE,因此在計算上獎勵在Sociology本身上的比例較小,而非特定的SBE的比例較大
以計畫的主題共現網路圖來分析,網路圖上節點代表該計畫內出現的各學科,節點大小表示相對應學科所占獎勵數量比例,節點間的連接線則代表兩個學科曾至少共同出現在一個獎勵,亦即它們之間曾有科際整合的記錄,線的粗細代表它們共同的獎勵數量比例。圖形上節點的數量表示對應計畫內學科的種類數量(variety),平均相連程度和網路密度則可以用來測量學科間的互動程度。例如,DRMS和Economics的網路上各有23個不同的學科,但是DRMS的平均相連程度和網路密度都比Economics來得大,分別是0.345 vs. 0.252和9.74 vs. 7.13,其原因是DRMS計畫內的獎勵大多由3到4個學科組成,而Economics的獎勵則只有2個學科。

2015年3月24日 星期二

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

Yan, E. (2014). Research dynamics: Measuring the continuity and popularity of research topics. Journal of Informetrics, 8(1), 98-110.

由於發現新的物種、疾病與社交模式,產生新的研究主題與專業 (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012),經過一段時間後,相關的研究社群會成長或是規模改變,有些主題仍然持續,但有些則是消失 (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010)。已有許多研究利用書目資料來確認研究的專業,例如Kessler (1963)的論文書目耦合網路(bibliographic coupling networks)、Small (1973)的論文共被引網路(paper co-citation networks)、White 與 McCain (1998)的作者共被引網路(author co-citation networks)以及White (2003)的尋路網路 (pathfinder networks),Callon、Courtial 與 Laville (1991)、Ding、Chowdhury 與 Foo (2000)、Milojevic、Sugimoto、Yan 與 Ding (2011)則是使用詞語共現網路 (co-word networks)。這些研究各自在不同研究層次確認研究主題,例如論文層次有 Chen (2004, 2006)、 Kessler (1963)和 Small (1973),作者層次有 Clauset, Newman, & Moore (2004)、White & McCain (1998)和 White (2003),期刊層次如 Glänzel & Schubert (2003)、 Leydesdorff & Vaughan (2006),以及領域層次有 Janssens, Zhang, Moor, & Glänzel (2009)、Rafols & Leydesdorff (2009)、Zhang, Liu, Janssens, Liang, & Glänzel (2010)。較低的研究實體層級,如論文與作者,研究可以從領域內發現其他的主題或專業;但在期刊或領域等較高的層次,通常從更完整的資料中確認出次領域。確認主題的方法則有因素分析(factor analysis)和多維尺度(multidimensional scaling)等傳統的群集技術以及連結線中心性(edge betweenness)、群組性(modularity)和混合群集(hybrid clustering)等較新技術的應用。本研究(Yan, 2014)則是利用主題模型(topic model)確認研究主題,並提出主題延續性(topic continuity)及主題普遍性(topic popularity)等兩項動態特性來分析研究主題。應用主題模型技術考察主題動態的方法,包括事後分析(post hoc analysis)(例如: Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008)、分段法(segmented approaches) (例如:Bolelli, Ertekin,Zhou, & Giles, 2009)以及連續時間模型(continuous-time model) (Wang & McCallum, 2006)等。本研究採用事後分析,利用文件中各主題的機率分布評估主題存在的機率。

本研究針對每一年分別產生一個主題模型,對於每一個主題找出後一年最有可能的主題,評估兩個主題相似的方式是利用改良自Kullback-Leibler差異 (Kullback-Leibler divergence, KLD)的Jensen–Shannon差異 (Jensen–Shannon divergence, JSD),兩個模型P和Q的JSD計算方式為JSD(P||Q) = 1/2KLD(P||M)+1/2KLD(Q||M),KLD是兩個模型Kullback-Leibler差異,M=1/2(P+Q)。每一個主題後一年最有可能的主題是擁有最小JSD的主題,其分數JSD稱為JJSDS,運用JJSDS的變化趨勢計算主題的連續性,然後以z score進行標準化。
評估各主題的普遍性則是計算它們在該年度文件上平均的機率值,並以z score進行標準化,愈大的機率值表示該主題在當年度愈普遍,然後分析主題普遍性的變化趨勢,。

本論文的研究資料為2001到2011年的圖書資訊學(library and information science)出版品,包括期刊論文、書評及研討會論文等,採用論文的題名做為分析資料,共27,796 篇論文。每一年的主題數目都設為20。結果顯示在網路資訊檢索(web information retrieval)、引用及書目計量學(citation and bibliometrics)、系統及技術(system and technology)、健康科學(health science)等主題有較高的平均普遍性;h指標(h-index)、線上社群(online communities)、資料保存(data preservation)、社群媒體(social media)和網站分析(web analysis)等則是圖書資訊學裡愈來愈普遍的主題。研究結果的主題與過去的研究相符合,但這篇論文的貢獻在於對於研究主題的動態進行分析。

Dynamic development is an intrinsic characteristic of research topics. To study this, this paper proposes two sets of topic attributes to examine topic dynamic characteristics: topic continuity and topic popularity.

Topic continuity comprises six attributes: steady, concentrating, diluting, sporadic, transforming, and emerging topics; topic popularity comprises three attributes: rising, declining, and fluctuating topics.

These attributes are applied to a data set on library and information science publications during the past 11 years (2001–2011).

Results show that topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity; topics on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular in library and information science.

Dynamics is a constant theme in scientific explorations. Research communities may grow or change in size; new species, diseases, or societal patterns may be discovered; and new research topics and specialties may be introduced (Li et al., 2010; Yan, Ding, Milojevic, & Sugimoto, 2012). Over time, some topics are continuously investigated while others appear or disappear (Griffiths & Steyvers, 2004; Upham & Small, 2010; Shi, Nallapati, Leskovec, McFarland, & Jurafsky, 2010). Therefore, it is of great importance to examine research dynamics to understand the evolving cognitive structures of research domains.

Pioneering studies of paper bibliographic coupling networks (Kessler, 1963), paper co-citation networks (Small, 1973), author co-citation networks (White & McCain, 1998), pathfinder networks (White, 2003) and co-word networks (e.g., Callon, Courtial, & Laville, 1991; Ding, Chowdhury, & Foo, 2000; Milojevic, ´ Sugimoto, Yan, & Ding, 2011) were capable of identifying research specialties from bibliographic data effectively.

However, findings from these studies remained largely static and thus only yielded fixed perspectives on the cognitive structure of research domains.

To examine research dynamics,this study uses a topic modeling technique and proposes two sets of topic attributes–topic continuity and topic popularity.

• How to use topic modeling techniques to study research dynamics?
• What quantitative measurements can be used to describe topic dynamics?
• What topics are present in library and information science? What are their dynamic characteristics?

This subsection reviews the network-based approaches of identifying research topics and specialties. These approaches have been applied to several research levels, including the paper-level (e.g., Chen, 2004, 2006; Kessler, 1963; Small, 1973), the author-level (e.g., Clauset, Newman, & Moore, 2004; White & McCain, 1998; White, 2003), the journal-level (e.g., Glänzel & Schubert, 2003; Leydesdorff & Vaughan, 2006), and the field-level (e.g., Janssens, Zhang, Moor, & Glänzel, 2009; Rafols & Leydesdorff, 2009; Zhang, Liu, Janssens, Liang, & Glänzel, 2010).

Most above-mentioned work used co-occurrence networks as the research instrument.

Analyses on lower level research entities, such as papers and authors, usually identified topics and specialties from small but well-defined research fields; whereas analyses on higher level research entities, such as journals and fields, attempted to identify subfields and subdomains from more comprehensive data sets.

Both classic clustering techniques (e.g., factor analysis and multidimensional scaling) as well as modern techniques (e.g., edge betweenness, modularity, and hybrid clustering) have been applied.

Recently, studies have attempted to add dynamic analyses by utilizing multiple time intervals.

Several approaches on slicing time intervals are available: intervals that have the same amount of references (e.g., Radicchi et al., 2009), intervals that have the same number of publications (e.g., Sugimoto, Li, Russell, Finlay, & Ding, 2011; Yan & Sugimoto, 2011), same-length intervals (e.g., Åström, 2007; Milojevic´ et al., 2011), and accumulative intervals (e.g., Barabási et al., 2002; Yan & Ding, 2009).

These studies laid valuable methodological basis for dynamic analyses of cognitive structures of research fields; however, networks of different time frames were largely analyzed distinctively and a more integrated examination was lacking.

In the meantime, empirically, network-based clustering results may require domain expertise to effectively interpret obtained results.

Topic modeling techniques use probabilistic models to assign papers, journals, or authors to clusters. A topic can be defined as a probability distribution over terms in a vocabulary (Blei & Lafferty, 2007). Latent Dirichlet Allocation (LDA) model, a classic topic model, was proposed by Blei et al. (2003). The model predicates that words for each paper are derived from a mixture of topics and each topic follows a multinomial distribution.

One recent update of the LDA model is the supervised LDA model. It makes the analyses of multi-labeled corpora (e.g., tags from delicious.com and various classifications) possible. Blei and McAuliffe’s (2010) version of supervised LDA can successfully address this challenge, but a document can only be assigned with one label.

Ramage, Hall, Nallapati, and Manning (2009) offered an approach which enabled the multi-label assignment. Their supervised labeled LDA (L-LDA) associated one label with one topic and allowed the model to learn word-label relations.

Through topic modeling techniques, topic dynamics has been examined mainly through the following approaches: post hoc analysis (e.g., Griffiths & Steyvers, 2004; Hall, Jurafsky, & Manning, 2008), segmented approaches (e.g., Bolelli, Ertekin,Zhou, & Giles, 2009), and continuous-time model (Wang & McCallum, 2006).

Post hoc analysis uses topic-document probability distributions to evaluate the presence of identified topics.

Segmented approaches build the dynamic component in the probabilistic model. It assumes that the state of topics at a single time point is independent from all other time points and divides document corpora into segments that have contingent time stamps (Bolelli et al., 2009).

Continuous-time model is a non-Markov model proposed by Wang and McCallum (2006), where they found the non-Markov model provides better prediction and more interpretable topical trends.

In this study, a post hoc dynamic analysis using the ACT model is selected because of its marked performance (Tang et al., 2008) as well as its advanced input and output support.

Topic dynamics is calculated through the Author-Conference-Topic (ACT) model (Tang et al., 2008).

Specifically, i is the topic distribution for document i. Mean ( ¯), therefore, is a direct quantitative measurement to assess topic popularity: the higher the ¯, the more visible the topic, and thus the more popular that topic is (Griffiths & Steyvers, 2004).

Because the data set spans 11 years, 11 independent ACT models were run, one for each year of the data set based on year of publication.

The Jensen–Shannon divergence (JSD) was used as the similarity measurement to quantify the topic similarity between different word-topic distributions. ... JSD is a symmetrized and smoothed version of the Kullback–Leibler divergence (KLD). ... As a divergence measure, the smaller the JSD, the higher the similarity is.

In order to track the same topic from two adjacent time intervals, the minimum value for each row of a JSD matrix was used, referred to as the joint JSD score (JJSDS): MIN(JSD Matrix(i,j)), for j = 1:n. ... Applying the same approach to each pair of adjacent time slices, for each topic, an array of JJSDS can be obtained.

The attributes of steady, concentrating, and diluting topics focus on the overall topical characteristics whereas the attributes of sporadic, transforming, and emerging topics focus on the topical characteristics of a specified time frame. Therefore, these attributes are not mutual exclusive, suggesting that a topic can be a concentrating topic overall, and in the meantime, related topics were added and thus qualifying it for a transforming topic.

The data set contains publications of all journals indexed in the 2011 version of the Journal Citation Report in the Information Science & Library Science subject category. Articles, proceeding papers, and review articles published within these journals from 2001 to 2011 were downloaded for analysis (downloading time: October 2012). Stop words were then removed from publications’ titles. Publications without titles, authors, or journal names were removed from the data set. The final data set comprised 27,796 papers.

The number of topics is set at 20: this number considers the size of the paper corpus as well as previous empirical studies on the cognitive structure of library and information science (e.g., Milojevic´ et al., 2011; Sugimoto et al., 2011; White & McCain, 1998; Zhao & Strotmann, 2008). For reasons of consistency, the same number of topics was identified for each year of the data set.

In this subsection, we first present histograms made from values in Jensen–Shannon divergence (JSD) matrices (Fig. 4). These histograms provide a direct perception on how research topics in library and information science are related as measured by JSD. This subsection then introduces all 20 topics in each year from 2001 to 2011 as well as how topic continuity and popularity attributes are applied to these topics (Fig. 5).

Fig. 4 uses histograms to visualize JSD values for each pair of adjacent years. Because there are 20 topics for each year, the number of data points in each histogram is 400 (20 × 20). This number is 4000 for the histogram in the lower right section of Fig. 4, as it uses JSD values for all pairs of adjacent years.

This study finds that in library and information science, research topics on “web information retrieval”, “citation and bibliometrics”, “system and technology”, and “health science” have the highest average popularity over the past decade (from 2001 to 2011).

Research on “h-index”, “online communities”, “data preservation”, “social media”, and “web analysis” are increasingly becoming popular topics.

Overall, findings of this study are consistent with previous studies using co-word, co-citation, and topic modeling techniques.

For instance, a co-word study by Milojevic´ and colleagues (2011) has found that title terms “citation”, “impact factor”, and “web” have a rising usage from 1989 to 2008.

Other related dynamic studies that cover the target time frame of the current study (2001–2011) include Åström’s (2007) study on examining library and information science research front, where the study found that webometrics and information-seeking and retrieval have become dominating research areas between 2000 and 2004.

This finding has been verified by Klavans and Boyack (2011) where the authors used the global map (i.e., the map of science) to enhance to accuracy of local maps (i.e., the contextual map of information science). They identified five core areas in information science, including information-seeking behavior, computer-enhanced retrieval, scientometrics, co-citation analysis, and citation behavior.

Besides the contextual analysis of information science, structural analysis has also been achieved from a time-series empowered author co-citation and document co-citation analysis (Chen, Ibekwe-SanJuan, & Hou, 2010). Through the application of a series of structural metrics such as centrality measures, modularity and silhouette, a clear cognitive structure of information science was attained in that the research areas of interactive information retrieval, academic web, information retrieval, citation behavior, and h-index have gained a particular popularity from 1996 to 2008.

In addition to journal publications, Sugimoto and colleagues (2011) applied a LDA model to library and information science dissertations and demonstrated dissertations as an important communicative genre. Their study indicated that between 2000 and 2009, internet and information retrieval related topics were the central dissertation research themes.

The contribution of the current study is that it proposes two sets of quantitative topic attributes. These attributes have streamlined the dynamic analysis of research topics and specialties and have further complemented co-occurrence-based studies.

This paper has identified dynamic characteristics of topics in library and information science; however, limited information can be told about the mechanisms that resulted in such characteristics. That being said, the study is unable to pinpoint, for instance, whether the growing popularity of network and citation studies is the result of a growing research community, a drive by the commercial market, a stimulus from funding agencies, or a combination of these or other unlisted factors.

Popular topics may be associated with research communities that are expanding in size and/or tend to have higher productivity. Conversely, less popular topics may be associated with communities that are shrinking and/or have a reduced productivity. Topic continuity and popularity attributes reflect research specialties’ development in scientific communities, which is further guided by science policies and the attention of the general public.

In informetrics, studies have mainly focused on analyzing the performance and the social and cognitive implications of several types of research entities, including papers, authors, institutions, journals, and fields. Authors and institutions are typically used to examine social relations in academia; while journals and fields are predominantly used to investigate the cognitive structure of research domains.

Topic analysis can precisely provide a more refined assessment by clustering research papers based on certain probability distributions. Because of such quantitative results, a more integrated dynamic cognitive analysis is thus possible, as exemplified through the current study.

Topic analysis will be further developed by overlaying topics with author communities to explore the interwoven relationships between research topics and research communities (e.g., Yan et al., 2012); by overlaying topics with funding data to investigate the “lead-lag” relationship between funding support and productivity (e.g., Shi et al., 2010); by applying topic models to different genres to study research immediacy (e.g., Ding et al., 2013); and by overlaying topics with citation data to examine the relationships between topics and impact.

2014年8月10日 星期日

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers. Scientometrics, 94(3), 1161-1173.

Ni, C., Sugimoto, C. R., & Cronin, B. (2013). Visualizing and comparing four facets of scholarly communication: producers, artifacts, concepts, and gatekeepers.Scientometrics, 94(3), 1161-1173.

network analysis

本研究以發表場域-作者-耦合(Venue-Author-Coupling,VAC)、期刊共被引分析(journal co-citation analysis)、主題分析(topic analysis)和連結編輯委員會成員(interlocking editorial board membership)等四個面向分析資訊科學與圖書館學的期刊網絡。這個研究分析的期刊範圍為2008年 JCR (Journal Citation Report)資訊科學與圖書館學分類的58種期刊,在2005到2009年間的出版資料。分析資料的相關數據如Table 1:


本研究利用VAC代表期刊的生產者(producers)的相似性,根據每一對期刊間相同的作者數量測量它們的接近程度,其原理建立在作者會選擇主題或社會性相似(thematically or socially similar)的期刊發表。期刊共被引分析(McCain, 1991)計算每一對期刊被共同引用的次數,本研究用來測量作品(artifacts)間的相似程度。本研究以修改自LDA模型(Blei et al. 2003)的ACT(Author-Conference-Topic)模型(Tang et al., 2008)透過關鍵詞(keywords)在主題上的分布以及主題在作者及發表場域(期刊)上的分布,本研究以餘弦(cosine)測量評估期刊之間的相似程度。連結編輯委員會成員則是編輯委員會上的共同成員數測量期刊間的相似程度。兩種期刊間共同的成員愈多,代表這兩種期刊在認知上或是社會性上愈相似。

根據上面的四種期刊間的相似程度所得到的結果,除了進行階層式集群分析(hierarchical cluster analysis)之外,也用來建立網絡,以Kamada-Kawaii 法呈現網絡的型態。分析得到的四種網絡並且以二次指派程序(Quadratic Assignment Procedure) (Lawler 1963)比較網絡之間可能的相關性(correlation)。

VAC方法得到的期刊網絡如下
四個集群分別為MIS(黃)、IS(藍)、LS(綠)以及專門性期刊(紅)。其中的MIS期刊集群與其他的集群相當分離。IS與LS距離較近。相較於其他三個期刊集群,專門性期刊彼此間的連結較弱。
期刊共被引分析所得的網絡如下:

主題模型產生的五個主題如Table 2
五個主題在網絡上的分布如下圖
MIS(黃)仍然與其他集群較為分離,但與健康和傳播(communication)等專門性期刊的距離較近。IS(藍)和LS(粉紅)的位置與VAC和期刊共被引分析的網絡上有所不同。圖書館服務與實務(綠)與專門性期刊和LS很接近。

在利用連結編輯委員會成員的期刊網絡上,有10種期刊沒有和其它期刊有共同編輯委員。其餘的集群分為四群。以傳播研究相關的期刊是新增加的集群(綠)。

四個網絡的QAP結果如Table 3

總結以上,在JCR的資訊科學與圖書館學分類下約略可以將期刊分為四個集群:MIS、IS、LS和傳播相關的期刊。MIS相較來說較為獨立。另外,QAP的結果可以看到編輯委員會成員的結果與期刊共被引分析有很高的相關性,其原因可能是由於擔任編輯委員的研究人員往往有較好的學術成就,被引用的機會較高。編輯委員會成員與VAC有較高的相關性,其原因也可能是編輯委員有較高的生產力。運用多種面向的分析可以較全面地了解整個學術傳播網絡。

Fifty-eight journals from the Information Science and Library Science category in the 2008 Journal Citation Report were studied and the network proximity of these journals based on Venue-Author-Coupling (producer), journal co-citation analysis (artifact), topic analysis (concept) and interlocking editorial board membership (gatekeeper) was measured. The resulting networks were examined for potential correlation using the Quadratic Assignment Procedure.

The VAC approach is used to represent the producers in this dataset. This approach measures journal proximity based on the number of authors shared by each journal pair. The VAC approach is based on the idea that an author’s choice of publication venue reflects similarity judgments authors are likely to choose venues that are thematically or socially similar.

Artifacts are measured by means of journal co-citation. This measure, introduced by McCain (1991), refers to the appearance of two journals in the same reference list of an article. The more frequently two journals appear in the same reference lists, the greater the similarity between the two journals. The journal co-citation approach measures journal proximity by the frequency with which each journal pair is co-cited by the same articles.

Topic modeling is used to capture concepts. ... The technique adopted here, the author-conference-topic (ACT) model (Tang et al., 2008), extends the LDA model by considering the author and publishing venue of the articles. LDA was developed originally as a topic modeling technique concerning the probability distribution of keywords for topics, and is particularly helpful with the ‘‘classification, novelty detection, summarization, and similarity and relevance judgment’’ of large-scale data (Blei et al. 2003, p. 993). ... This model extends the idea of LDA by taking into account the authors and publishing venues, and estimates not only the distribution of words on topics, but also the distribution of authors and venues on the topics modeled. ... Here, the outcome of the ACT model is the probability distribution of each author and each journal over topics, and the journal proximity is calculated using the cosine similarity of the journals.

The interlocking editorship approach, employed by Ni and Ding (2010), measures journal proximity based on common editorial board membership. The number of editorial board members that two journals share can be viewed as an indicator of journal similarity. ... Thus, it can be expected that if two journals have scholars in common on their editorial boards, these two journals have some degree of similarity, either cognitively or socially.

The journals were clustered using a hierarchical clustering technique with squared Euclidean distance and Ward’s method. Each journal clustering was displayed as a network (Kamada-Kawaii layout); each node (journal) was colored according to the hierarchical clustering result with the size of a
node proportional to its centrality (either degree or closeness).

Additionally, a comparison of journal proximity results was conducted using the Quadratic Assignment Procedure (QAP). QAP is commonly used in social network analysis as a means of investigating correlations between two networks. ... (Lawler 1963).

2014年6月13日 星期五

Bache, K., Newman, D., & Smyth, P. (2013, August). Text-based measures of document diversity. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining (pp. 23-31). ACM.

Bache, K., Newman, D., & Smyth, P. (2013, August). Text-based measures of document diversity. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining (pp. 23-31). ACM.

Scientometrics

本研究提出一個利用文本內容測量多樣性的架構,其原理是先以文件語料庫訓練主題模型,然後根據主題在文件上的共現情形,計算主題之間的距離,最後再測量文件的多樣性。這個架構的優點是只需要以文件語料庫的文本資料做為輸入,完全是資料驅動(data-driven)的方法;產生人類可讀的(human-readable)結果;並且能夠廣泛地擴大到作者、學系和期刊等應用。
計算某一個群體的多樣性(diversity)程度,目前已經廣被各種學科關注,例如生態學(ecology) [9]、遺傳學(genetics) [12]、語言學 (linguistics) [8]和社會學(sociology) [5]。在評估群體的多樣性時,通常假定群體內的個體可分為T個類別,每個類別的比例為pi:
Stirling [22]提出多樣性應包含三個面向:類別數目(variety)、比例之間的平衡(balance)以及類別間的差異(disparity)。
1. 群體內包含的類別數目:是相對簡單的多樣性測量方式,也就是π裡非零比例的數量。

2. 比例之間是否相對平衡,常用的測量方式包括Shannon的熵(entropy),或是變異量

3. 群體上呈現的類別彼此間的差異。
Stirling [22]並說明這三個面向都是多樣性量化的必要條件,但不是充分條件,而由Rao [18]提出的公式可以將三個面向整合起來,

此處pi和pj分別類別i和j在群體中的比例,δ (i, j) 是類別i和j的距離,Δ是 TXT的距離矩陣,t表示矩陣轉置運算。

Rafols and Porter [14]利用引用文獻的ISI 期刊主題分類(journal subject categorizations)分析6個特定的主題分類在1975到2005年間的跨學科情形。在他們的研究中,是參考文獻發表期刊的主題類別(subject category) i,δ (i, j)定義為1減去主題類別i和j的引用計數向量(citation count vector)的cosine距離。結果發現雖然引用文獻數量與共同作者數有明顯的增加,跨學科性的程度以緩慢的速度增加。這種方法仰賴事先定義的分類而有限制 (Rafols and Porter [15]),因為主題分類可能隨時間改變,而無法反應當時的學科界線。當分析的資料沒有適合的分類方式時,此時也會受到限制。此外,引用資料能否準確反應科學文獻的內容也頗受爭議。

In this context we can define Rao's diversity measure for each document d as

本研究提出一個文本為基礎的研究架構,以文件的內容來測量它的多樣性程度。此一方法利用Latent Dirichlet Allocation (LDA)的主題模型(topic model)方法從文件的語料庫(corpus)推論出T個主題,最後產生出一個DXT的矩陣,D是文件的數量,矩陣上的元素ndj表示文件d上的詞語指定給主題j的比例。利用此一矩陣以Cosine距離計算每一對主題間的距離矩陣,然後結合這些距離和文件上的主題分布,利用Rao [18]提出的公式估算文件的多樣性。
此一方法是完全資料驅動(data driven),產生容易解讀的結果,並且能夠將其推廣應用於作者、學術部門以及期刊的多樣性測量。


In this paper we present a text-based framework for quantifying how diverse a document is in terms of its content.

The proposed approach learns a topic model over a corpus of documents, and computes a distance matrix between pairs of topics using measures such as topic co-occurrence. These pairwise distance measures are then combined with the distribution of topics within a document to estimate each document's diversity relative to the rest of the corpus.

The method provides several advantages over existing methods. It is fully data-driven, requiring only the text from a corpus of documents as input, it produces human-readable explanations, and it can be generalized to score diversity of other entities such as authors, academic departments, or journals.

The quantification of diversity has been widely studied in areas such as ecology [9], genetics [12], linguistics [8], and sociology [5].

The typical context is where one wishes to measure the diversity of a population, where a population consists of a set of individual elements that have been categorized into T types (such as species), with proportions 

A relatively simple measure of diversity is variety, how many different species are present in a population, or the number of non-zero proportions in π.

One can alternatively measure diversity as a function of the relative balance among the proportions (also referred to as `evenness' in ecology [13] or `concentration' in economics [4]), using measures such as Shannon entropy or variance-based quantities such as 

From a more general perspective, Stirling [22] proposed that there are three distinct aspects to diversity: variety, balance, and disparity.

Disparity is the extent to which the categories that are present are different from each other, based for example on distance within a known taxonomy [21].

Stirling argued that each of these three properties is a necessary (but non-sufficient) component in any quantitative characterization of diversity, arriving at a relatively simple mathematical formulation for diversity, a formulation originally proposed in earlier work by Rao [18]:


where pi, pj are the proportions of category i and j in the population, δ (i, j) is the distance between categories i and j, Δ is a TXT matrix of such distances, and t is the transpose of the TX1 vector of proportions .

The contribution of this present paper is to investigate diversity in the context of text documents, using Rao's measure a starting point.

In particular, we will use words as elements, topics as word categories, and documents as collections (or "populations") of words. Specifically, we address the following task: given a corpus of documents, assign a diversity score to each document, where this diversity score can be used to rank documents from most to least diverse.

Indeed, diversity as defined via co-citation counts is the most widely-used approach to quantify interdisciplinarity in practice, based on the notion that disciplines that are co-cited more often by the same article are "closer" than disciplines that are less frequently co-cited.

Journal subject categories are typically used to capture the notion of a discipline, typically using the manually-defined 244 ISI subject categories from Thomson Reuters, with articles being assigned to a subject category associated with the journal the article is published in (e.g., [15, 14, 17, 23]).

Rafols and Porter [14] used journal subject categorizations of citations to analyze how interdisciplinarity has changed between 1975 and 2005 for six specific subject-categories. They concluded that although the number of citations and co-authors per paper was increasing significantly over time, the degree of interdisciplinarity was increasing at a much slower rate, as reflected by citation patterns between subject categories. As a component in their analysis, Rafols and Porter used Rao's diversity index based on a count matrix of D documents by T categories derived from citations: pi was the proportion of citations made by an article to other articles that were published in journals belonging to subject category i, and δ (i, j) was defined as 1 minus the cosine distance between citation count vectors (across documents) of subject categories i and j.

Our work differs from this earlier work and related threads in scientometrics in two specific ways. First, in our approach the categories and distances, δ (i, j), are learned directly from the text content, rather than being based on manually predefined schema such as the ISI subject categories. ...  The second major difference in our approach is our use of word counts rather than citation counts (which are the basis of most prior work in scientometrics on quantifying interdisciplinarity). 

There are obvious limitations to relying on pre-defined taxonomies, as pointed out by Rafols and Porter [15]. Subject categories can change over time and no longer necessarily reflect current disciplinary boundaries.

In addition, in some contexts such as analysis of proposals and grants, there may be very limited or no categorizations available. For analysis of narrow domains (say the field of data mining and machine learning) existing categorization schemes may be too coarse-grained to be useful. In this context, a corpus-driven approach to learning the categories, such as the topic-based method we describe here, is a useful alternative, and in some cases may be the only option.

We expect that using text content will complement citation-based approaches, as both words and citations carry useful signal. There has long been debate over whether citations accurately reflect the content of a scientific article [2, 1]-- arguably the words in an article may provide a more accurate reflection of the author's intentions than the citations the author uses.

Another field which is related to our current work is that of outlier detection. If we consider documents as being represented by T-dimensional vectors of counts, then one approach to quantifying diversity is to look for documents that are outliers in this T-dimensional space, using a multivariate outlier detection algorithm. ... Equivalently, since the pi are the components of a probability vector in a T - 1 dimensional simplex, we can think of high diversity documents as points that lie in the interior of the simplex (in at least 2 of the dimensions) rather than at the edge.

We use the Latent Dirichlet Allocation (LDA) topic model with collapsed Gibbs sampling to learn T topics for the D documents in the corpus [7]. A single iteration of the collapsed Gibbs sampler consists of iterating through the word tokens in the corpus, sequentially sampling topic assignments for each word token in each document while keeping all other topic-word assignments fixed. Using the topic-word assignments from the final iteration of the Gibbs sampler , we create a DX T document-topic count matrix with entries ndj corresponding to the number of word tokens in document d that are assigned to topic j.

In this context we can define Rao's diversity measure for each document d as

where P(j|d) is the proportion of word tokens in document d that are assigned to topic j and δ (i, j) is a measure of the distance between topic i and topic j. Note that δ (i, j) is constant across all documents, and P(i|d) and P(j|d) vary from document to document.



We presented an approach for quantifying the diversity of individual documents in a corpus based on their text content. Empirical results illustrated the effectiveness of the method on multiple large corpora.

This text-based approach for assigning diversity scores has several potential advantages over previous alternatives, such as methods that define diversity based on citations categorized into predefined journal subject categories. The text-based approach is more data-driven, performing the equivalent of learning journal categories by learning topics from text, and can be run on any collection of text documents, even without a prior categorization scheme.

In addition, it produces human-readable explanations and can be easily generalized to score the diversity of other entities such as authors, departments, or journals (e.g., by aggregating counts across such entities).

2013年12月7日 星期六

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

本研究提出一個TTR-LDA-社群模型,這個模型以推論機制(inference mechanism)結合LDA(Latent Dirichlet Allocation)模型和Girvan-Newman社群偵測(community detection)演算法提供在網路資料上偵測社群並對這些社群進行主題探勘(topic mining)的功能,並且進而了解在社群上的主題隨時間推移的變化,處理的架構如下圖所示

本研究利用Delicious社會標籤系統(social tagging system)上從2005到2008年的資料進行研究。在社群偵測部分,首先建立網絡:根據使用者標籤的資源數量,選取前50000位標籤資源最多的使用者;然後對他們標籤的網頁進行統計,從其中選取10000個被最多使用者標籤的網頁。接著在上述的10000個網頁中,如果有這50000位使用者之間有任何兩位曾經標籤過相同的網頁,便在這兩位使用者之間產生一個連結。以50000個使用者為節點,同時以他們之間的連結為連結線,便可以建立一個共同書籤網絡(co-bookmark network)。並且為了研究網絡上社群結構的變化,並將整個期間的資料分為三個時段:分別為2005-2006、2007與2008年,相關的統計數據如下表:

本研究利用Girvan-Newman演算法找出標籤者(Tagger)社群,使得標籤者與同一社群內的其他標籤者比社群外的標籤者有較強的關係。這個演算法重複移去當時網絡上中介性(betweenness)最大的連結線,產生各種可能的網路劃分(network partition),測量每一種劃分下的群組性(modularity),也就是實際上社群內的成員彼此間的連結線數量與相同連結度的情況但隨機產生連結線的數量的差,群組性最大的劃分便是輸出結果。

另一方面,本研究利用TTR-LDA模型找出每個標籤者的主題分布以及主題內具有代表性的標籤,TTR-LDA模型修改自ACT(author-conference-topic)模式[11][12],是一個由標籤者做為第一層、標籤與資源為第三層、主題則為第二層,所構成的三層貝氏模型(three-layer Bayesian model)。

整合Girvan-Newman演算法和TTR-LDA模型的方法是以社群為單位,將社群內所有標籤者的主題分布進行平均做為該社群的主題分布,根據主題分布,選出機率值較大的主題做為社群的代表。比較不同時段社群共同的代表標籤衡量它們的相似性。

社群偵測的結果發現前五個最大的社群在四年裡占了絕大多數的比率,而且這個比率逐年增加。主題探勘的部分則測量TTR-LDA模型在不同主題數量上的複雜度(perplexity),發現150個主題時有最低的複雜度。因此,以下的研究便針對150個主題在前五個最大社群上的分布進行探討,計算它們的傳導性(conductance)與模組性。就社群模組性而言,最近一個時段(2008年)的結果比前三個時段還要高,其原因可能是因為經過一段時間後,社群的結構逐漸成熟,因此後期比前期更能產生較佳的社群。另外,將最後一個時段再細分為四個較小的時段則發現,較小時段的社群模組性比整年的結果來得高,本研究認為造成這種現象的原因可能是由於在不同的時段,大部分標籤者的書籤行為集中在不同的領域;當那些時段合併起來的時候,會展現標籤者在多個領域的興趣,使得社群內的叢集(clustering)特性較弱。

結果並可以發現前二十個主題都出現在不同時段的前五個最大社群裡,每個社群至少包括一個前十名的主題。此外,比較LDA、TTR-LDA和TTR-LDA-社群等三種模型在資源和標籤上的預測力,在回收率(recall rate)、精確率(precision)和F1指標上以TTR-LDA-社群為最佳。

In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism.

The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008.

Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time.

From a research perspective, these real-world networks display unique properties from the classical random graph model [3] in that most real word networks exhibit three common properties: the small-world property, power-law degree distribution and a high clustering coefficient or transitivity (indicating community structure) [7][8][9].

Thus, an important task in network analysis is to detect communities and explore their features, which can improve community-supporting services at the community-level in the context of a social tagging system.

Many studies in various disciplines have been devoted to community detection; however, few of them have systematically and quantitatively studied the profiles of those detected communities.

In this paper, we propose a TTR-LDA-Community model, which is an inferential combination of an extended LDA model and a betweenness-based community detection algorithm. It provides rich, systematic, and quantitative information about the profiles of detected communities.

In the context of social tagging systems, where multiple users are annotating resources, the resulting topics reflect a shared view of the document; and the tags of the topics reflect a common vocabulary.

Girvan and Newman extended the betweenness measure to edges and designed a clustering algorithm which gradually removes the edges with the highest betweenness value [4]. This algorithm has been improved through modularity; and the complexity is reduced from O(m2n) to O(mdlogn) where d is the depth of the dendrogram of the community structure [2].

Many studies provide various models and algorithms for topic mining and community detection; yet, few of them have integrated those models and algorithms, performed topic mining for detected communities, and analyzed how those identified topics change among communities over time.

The activity of social tagging consists of three major components: tag, tagger and resource. The experimental dataset contains all the triples of these three components and the time and date of their creation on Delicious from 2005 to 2008.

In data processing, all taggers were ranked by the number of resources they have bookmarked and the top 50,000 taggers were selected as the sample of taggers.

These taggers bookmarked a total of 354,522 web pages, which were sorted by the number of taggers who bookmarked them. The top 10,000 resources were selected as the sample of web pages, associated with which a dominant majority of tagging activities occurred.

Thus a co-bookmark network was built in which a connection between two users (within the sample of 50,000 taggers) is created if they bookmarked the same resources (within the sample of 10,000 web pages).

In addition, in order to observe the evolution of structure and motif of communities, the time span (2005-2008) was divided into three slices.


The model is illustrated in Figure 1. TTR-LDA is developed based on ACT model [11][12]. It is a three-layer Bayesian model with taggers tap in each post p as the first layer, tags t, and resource r as third layer and all the topics denoted as latent variable z as the middle layer.

The inference mechanism is used to infer the topic distribution over detected communities.

Each community includes a set of taggers, who have a stronger relationship with other taggers within the community than the taggers outside.

Based on the taggers’ information model, the probability distribution of each tagger over a set of topics is obtained by using the TTR-LDA model while the community structure of taggers is revealed by the community detection algorithm. The two sets of results are further integrated through an inference mechanism.




Results show that the number of users of the top five communities occupies a major proportion in the four years (2005-2008) and the proportion is increasing over time.

Perplexity is used to identify the number of topics [10], which arrives at the lowest point when the number of topics is 150. The interest model of each tagger in the top five largest communities is then built based on their topic distributions.

By using users’ interest models and the inference mechanism, a topic distribution of the largest community can be created (Figure 3). We can find that the topic distributions in a community are diverse because users’ relationships in that community are mainly based on their co-bookmark activities not the similarity of their interest model.

In order to observe the dynamic features of communities, we design an experiment as follows:
1) denote the five largest communities from each time slice in 2008 as community_i_t where t means the tth time slice in 2008 and i means the ith largest community in tth time slice;
2) compute the topic distribution for the five communities, which is stored as model_t_i_Topic(j), the
probability of jth topic in ith largest community in the tth time slice;
3) obtain the probability distribution of tags that are collected from all the posts generated during the specific time slice; the probability of one tag occurring in a topic shows the level of representativeness of the tag for that topic;
4) sort all the tags according to their probability value in each topic and select the 20 top ranked tags to represent the content of the topics; select the top 5 ranked topics to represent the theme of each community;
5) analyze the similarity between different communities from different time slice through computing how many tags are shared by the two different communities. More specifically, we compare current time slice with its previous time slice, for example, we compare community_i_t with community_j_t-1 (j=1, 2…5).

The size of communities along evolutionary lines fluctuates over time. For example, the size of the community about social networks in the 3rd time slice (community_1_3) is much larger (4,377) than that (521) in the 4th time slice (community_5_4).

Conductance (from multi-criterion scores) and modularity (from single criterion scores) are used to evaluate the quality of communities detected by the TTR-LDA-Community model [6].

The smaller the value of conductance is, the higher the granularity of a community is. Network community profile (NCP) is used to compute and display the value of conductance for communities [5].

Whiskers networks and rewired networks are adopted as two comparative aspects. Whiskers is defined as the maximal sub graphs that can be detached from the rest of the network by removing a single edge; and a rewired network is a random network that has the same nodes and the same degree distribution as the original network [5].

The conductance of communities of the rewired original network (blue line in the left figure), rewired random network (red dashed line in the left figure), the original whiskers network (blue line in the right figure), and the random whiskers network (red dashed line in the right figure) are calculated and shown in Figure 4.

In Figure 4, compared with the rewired network (left) and the rewired whiskers (right), 1) the original network displays a higher granularity of communities (a lower conductance value); 2) the value of conductance as the function of the size of communities in the original network and the original whiskers present a “V” shape, showing properties of a true large social networks [5]; 3) the original whiskers has the best community granularity (the lowest conductance) between size 10-100; and 4) the best community granularity of rewired original network is around 1000.

The modularity of communities in the four time slices of 2008 is better than that in 2005-2007. This is probably due to the fact that community structure grows mature gradually over time, creating better communities in later years than in earlier years.

Meanwhile, modularity of communities in the short-term (four sub periods in 2008) is larger than the long-term (2008). It can be explained that in different time periods, most taggers’ bookmarking activities are focused on different domains, so in a certain short-term time period, communities may be quite different from each other. However, when those time periods are merged together, the taggers show different interests in many domains; so the clustering feature within the communities becomes weaker.

Results show that the most popular topics are about bandslash fiction, fan fiction, and supernatural fiction (the top 3 popular topics). Communities with similar theme are ranked 3rd, 4th, and 5th in size; and the web resources with similar topics are ranked 500-600 of the top 1000 ranked resources in number of taggers associated with them.

The top 20 ranked topics in 1000 most popular resources can be found in 5 largest communities in different time periods. For each community, there exists at least one topic that is ranked top 10 in 1000 most popular resources (Table 3).

Topic distributions for each community are obtained respectively from LDA, TTR-LDA model, and TTR-LDACommunity model based on co-bookmark network in a given period (Oct. 2008–Dec. 2008). One resource and five tags are recommended for each post according to the results of three models separately.

The TTR-LDA and TTR-LDA-Community model show significant improvement for recommendation of tags and resources for post in terms of precision, recall and F1-Measure. TTR-LDA and TTR-LDA-Community have slightly improved performance for “tags for post”, while TTR-LDA-Community outperforms TTR-LDA on “resource for post”.


2013年12月6日 星期五

Rosen-Zvi, M., Griffiths, T., Steyvers, M., & Smyth, P. (2004, July). The author-topic model for authors and documents. In Proceedings of the 20th conference on Uncertainty in artificial intelligence (pp. 487-494). AUAI Press.

Rosen-Zvi, M., Griffiths, T., Steyvers, M., & Smyth, P. (2004, July). The author-topic model for authors and documents. In Proceedings of the 20th conference on Uncertainty in artificial intelligence (pp. 487-494). AUAI Press.

本研究提出作者-主題模型(author-topic model),這是一個擴充LDA (Latent Dirichlet Allocation; Blei, Ng, & Jordan, 2003) 並加入作者資訊的文件產生模型。和基本的LDA相同的是都用一組的主題混合(mixture)來代表每一筆文件,但這個模型具有能夠由文件的作者決定表現文件上不同主題的混合權重(mixture weights)的特性。為了進一步說明這個模型,本研究比較了基礎的LDA、作者模型以及作者-主題模型的文件產生過程。首先,LDA的文件產生過程包含三個步驟:首先由一個Dirichlet分布中取樣,產生每一筆文件的主題分布;其次在產生文件中的每一個詞語之前,先從上面的主題分布中選取一個主題;最後從選定的主題對應的詞語分布上取樣產生這個詞語。整個模型如下圖所示

在這個模型裡,ϕ表示主題分布的矩陣,利用一個由詞彙裡的V個詞語的多元常態分布來表示T個主題中的每一個主題,而這些多元常態分布是從一個Dirichlet分布Dirichlet(β)中獨立地抽取而形成的。θ是文件上T個主題的混合權重所組成的矩陣,每一個文件的主題混合權重都是由一個Dirichlet分布Dirichlet(α)中獨立地抽取而形成的。最後,每一個文件中的詞語有一個相對應的主題zz是從文件相對應的主題混合權重θ取樣產生,而這個詞語則是根據主題z所對應的主題分布ϕ 所產生。運用演算法推導文件的產生模型,估算ϕθ可以分別提供語料中具有的主題以及這些主題在各個文件上的權重等資訊,常用的演算法包括變異推論(variational inference; Blei et al., 2003)、期望值延遲(expectation propagation; Minka & Laerty, 2002)和Gibbs取樣(Gibbs sampling; Griffiths & Steyvers, 2004)等。

作者模型如下圖所示

這個模型假設某一篇文件是由一群作者ad共同撰寫,某一位作者都擁有一套本身習慣使用的詞語組合,這個詞語組合可以由詞語的機率分布ϕ來描述,從一個Dirichlet分布Dirichlet(β)中獨立地抽取而形成的。在產生文件時,文件上的每一個詞語由ad中依據均勻分布(uniform distribution)隨機選取某一位作者x所撰寫,而詞語產生的機率便是由作者x對應的詞語機率分布來決定。因此若能估算ϕ,便能提供有關作者的研究興趣方面的資訊,並且進而從作者撰寫的文件內容的相似性估計他們在研究主題上的相似性。然而這個模型的問題是興趣主題的估算僅受限於作者所撰寫過的文件內容上的詞語,無法進一步擴及相同主題但使用不同詞語的文件。

作者-主題模型整合上述兩個模型。這個模型如同作者模型假設文件上的每一個詞語由一群共同作者ad中依據均勻分布(uniform distribution)隨機選取某一位作者x撰寫,但作者-主題模型假定每位作者本身擁有一套主題組合θθ上的機率分布是從Dirichlet(α)中獨立地抽取而形成的。所以在產生每一個詞語之前,先根據選定的作者x對應的主題機率分布θ挑選一個主題z,然後再根據這個主題對應的詞語分布機率ϕ挑選出一個詞語w做為輸出,詞語分布機率ϕ的產生則是由Dirichlet(β)中獨立地抽取而成。作者-主題模型如下圖所示

本研究運用Gibbs取樣推估各種模型的參數,這種方法提供根據Dirichlet先驗機率(prior)獲得參數估計的簡單方法並且允許從許多後驗機率分布(posterior distribution)的局部最大值(local maxima)組合估計值。

首先是LDA模型,這個模型包含兩組未知的參數:θϕ,以及隱藏的變數(latent variables)也就是指定給每一個詞語的主題z。通過採用在z上的Gibbs取樣(Gilks, Richardson, & Spiegelhalter, 1996),可以構建一個Markov鏈(Markov chain),並且這個Markov鏈收斂後驗分布(posterior distribution),然後利用這些結果推斷θϕ(Griffths & Steyvers, 2004)。在Markov鏈接續狀態間的轉移是來自於重複從以所有其他變數為條件下的分布上抽取z的結果。如下面的式子

此處zi = j代表文件的第i個詞語指定為主題j的情形;wi = m則代表第i個詞語實際上是詞彙中的詞語m的情形,z-i表示不包括第i個詞語的所有詞語的主題指定情形。CWTmj是不包括目前的案例下,詞語m被指定為主題j的次數,CDTdj則是不包括目前的案例下,主題j出現在文件d的次數。然後運用上面的式子產生Markov鏈,從上取樣所有訓練文件內各詞語的指定主題,並估算θϕ

此處ϕmj是在主題j中使用詞語m的機率,θdj 則是主題j在文件d中的機率。

運用同樣的方式來推測作者模型的未知參數ϕ,這個模型的隱藏變數是文件中每一個詞語所指定的作者x。在Markov鏈接續狀態間的轉移是在以所有其他變數為條件下的分布上重複抽取x的結果

此處xi = k 代表將第i個詞語指定為作者k的情形, CWAmk則是不包括目前的案例下,詞語m被指定為作者k的次數。所以作者k使用詞語m的機率可以用下面的式子推估



在作者-主題模型方面,這個模型包括兩組隱藏變數zx,在其他已知與未知變數的條件下,針對某一個詞語wi分別指定作者xik 與主題zij的機率

此處z-i,x-i代表不包括第i個詞語的所有詞語的主題與作者指定情形。CATkj 是不包括目前的案例下,作者k被指定為主題j的次數。根據上面的式子,給定主題j,詞語m被選取的機率ϕmj與給定作者k,主題j被選取的機率分別可以用下面的式子進行估算。



本研究以兩個資料集NIPS與Citeseer進行實驗,對作者-主題模型所發現的每一個主題,以從對應的主題機率分布內選取機率較高的詞語為該主題的標示,同時並選取該主題機率較高的代表作者。結果發現標示詞語大多能夠具體地代表不同的主題,而作者也大多是該主題非常知名的作者。NIPS資料集共計有1740篇研究會論文,共2037位作者,論文中所用的詞彙包括13649種詞語,資料集實際出現的詞語共有2301375次;CiteSeer資料集共計有162489篇論文摘要,共85465位作者,摘要中所用的詞彙包括30799種詞語,資料集實際出現的詞語共有11685514次。此外,本研究並且利用複雜度(perplexity)比較三種模型在詞語預測的成效。複雜度的計算如下

複雜度愈低表示該模型在詞語預測成效愈好。結果發現作者模型因為受限於模型本身概化的能力(generalization)較弱,所得到的成效比其他兩種主題模型差。因為加入來自作者的資訊,作者-主題模型在少量的訓練資料下有LDA較好的成效;當訓練語料增加,LDA便有較低的複雜度,並且當模型的主題數更多時,LDA需要的訓練語料更少。最後,本研究利用對稱 KL差異度(symmetric KL-divergence)測量作者的研究主題相似性,

並且利用熵(entropy)測量作者的研究主題分布廣度。

In our results we used two text data sets consisting of technical papers|full papers from the NIPS conference and abstracts from CiteSeer (Lawrence, Giles, & Bollacker, 1999). ... This leads to a vocabulary size of V = 13,649 unique words in the NIPS data set and V = 30,799 unique words in the CiteSeer data set. Our collection of NIPS papers contains D = 1,740 papers with K = 2,037 authors and a total of 2,301,375 word tokens. Our collection of CiteSeer abstracts contains D = 162,489 abstracts with K = 85,465 authors and a total of 11,685,514 word tokens.

We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words. A document with multiple authors is modeled as a distribution over topics that is a mixture of the distributions associated with the authors.

We apply the model to a collection of 1,700 NIPS conference papers and 160,000 CiteSeer abstracts.

Recently, generative models for documents have begun to explore topic-based content representations, modeling each document as a mixture of probabilistic topics (e.g., Blei, Ng, & Jordan, 2003; Hofmann, 1999).

With an appropriate author model, we can establish which subjects an author writes about, which authors are likely to have written documents similar to an observed document, and which authors produce similar work.

This generative model represents each document with a mixture of topics, as in state-of-the-art approaches like Latent Dirichlet Allocation (Blei et al., 2003), and extends these approaches to author modeling by allowing the mixture weights for different topics to be determined by the authors of the document.

In LDA, the generation of a document collection is modeled as a three step process. First, for each document, a distribution over topics is sampled from a Dirichlet distribution. Second, for each word in the document, a single topic is chosen according to this distribution. Finally, each word is sampled from a multinomial distribution over words specific to the sampled topic.

This generative process corresponds to the hierarchical Bayesian model shown (using plate notation) in Figure 1(a).



In this model, ϕ denotes the matrix of topic distributions, with a multinomial distribution over V vocabulary items for each of T topics being drawn independently from a symmetric Dirichlet(β) prior. θ is the matrix of document-specific mixture weights for these T topics, each being drawn independently from a symmetric Dirichlet(α) prior. For each word, z denotes the topic responsible for generating that word, drawn from the θ distribution for that document, and w is the word itself, drawn from the topic distribution ϕ corresponding to z.

Estimating ϕ and θ provides information about the topics that participate in a corpus and the weights of those topics in each document respectively.

A variety of algorithms have been used to estimate these parameters, including variational inference (Blei et al., 2003), expectation propagation (Minka & Laerty, 2002), and Gibbs sampling (Griffiths & Steyvers, 2004).

Assume that a group of authors, ad, decide to write the document d. For each word in the document an author is chosen uniformly at random, and a word is chosen from a probability distribution over words that is specific to that author.

x indicates the author of a given word, chosen uniformly from the set of authors ad. Each author is associated with a probability distribution over words ϕ, generated from a symmetric Dirichlet(β) prior. Estimating ϕ provides information about the interests of authors, and can be used to answer queries about author similarity and authors who write on subjects similar to an observed document.

However, this author model does not provide any information about document content that goes beyond the words that appear in the document and the authors of the document.

As in the author model, x indicates the author responsible for a given word, chosen from ad. Each author is associated with a distribution over topics, θ, chosen from a symmetric Dirichlet(α) prior. The mixture weights corresponding to the chosen author are used to select a topic z, and a word is generated according to the distribution ϕ corresponding to that topic, drawn from a symmetric Dirichlet(β).

In this paper, we will use Gibbs sampling, as it provides a simple method for obtaining parameter estimates under Dirichlet priors and allows combination of estimates from several local maxima of the posterior distribution.

The LDA model has two sets of unknown parameters -- the D document distributions θ, and the T topic distributions ϕ -- as well as the latent variables corresponding to the assignments of individual words to topics z. By applying Gibbs sampling (see Gilks, Richardson, & Spiegelhalter, 1996), we construct a Markov chain that converges to the posterior distribution on z and then use the results to infer θ and ϕ (Griffths & Steyvers, 2004). The transition between successive states of the Markov chain results from repeatedly drawing z from its distribution conditioned on all other variables, summing out θ and ϕ using standard Dirichlet integrals:

where zi = j represents the assignments of the ith word in a document to topic j , wi = m represents the observation that the ith word is the mth word in the lexicon, and z-i represents all topic assignments not including the ith word. Furthermore, CWTmj is the number of times word m is assigned to topic j, not including the current instance, and CDTdj is the number of times topic j has occurred in document d, not including the current instance.

For any sample from this Markov chain, being an assignment of every word to a topic, we can estimate and using


where ϕmj is the probability of using word m in topic j, and θdj is the probability of topic j in document d. These values correspond to the predictive distributions over new words w and new topics z conditioned on w and z.

An analogous approach can be used to derive a Gibbs sampler for the author model.

where xi = k represents the assignments of the ith word in a document to author k and CWAmk is the number of times word m is assigned to author k.



In the author-topic model, we have two sets of latent variables: z and x. We draw each (zi, xi) pair as a block, conditioned on all other variables:

where zi = j and xi = k represent the assignments of the ith word in a document to topic j and author k respectively, wi = m represents the observation that the ith word is the mth word in the lexicon, and z-i,x-i represent all topic and author assignments not including the ith word, and CATkj is the number of times author k is assigned to topic j, not including the current instance.

Equation 4 is the conditional probability derived by marginalizing out the random variables ϕ (the probability of a word given a topic) and θ (the probability of a topic given an author).




In the examples considered here, we do not estimate the hyperparameters α and β instead the smoothing parameters are fixed at 50/T and 0.01 respectively.

We start the algorithm by assigning words to random topics and authors (from the set of authors on the document). Each iteration of the algorithm involves applying Equation 4 to every word token in the document collection, which leads to a time complexity that is of order of the total number of word tokens in the training data set multiplied by the number of topics, T (assuming that the number of authors on each document has negligible contribution to the complexity).

In our results we used two text data sets consisting of technical papers|full papers from the NIPS conference and abstracts from CiteSeer (Lawrence, Giles, & Bollacker, 1999). ... This leads to a vocabulary size of V = 13,649 unique words in the NIPS data set and V = 30,799 unique words in the CiteSeer data set. Our collection of NIPS papers contains D = 1,740 papers with K = 2,037 authors and a total of 2,301,375 word tokens. Our collection of CiteSeer abstracts contains D = 162,489 abstracts with K = 85,465 authors and a total of 11,685,514 word tokens.

Perplexity is a standard measure for estimating the performance of a probabilistic model. The perplexity of a set of test words, (wd, ad) for d ∈ Dtest, is defined as the exponential of the negative normalized predictive likelihood under the model,



Better generalization performance is indicated by a lower perplexity over a held-out document.

The author model is clearly poorer than either of the topic-based models, as illustrated by its high perplexity. Since a distribution over words has to be estimated for each author, fitting this model involves finding the values of a large number of parameters, limiting its generalization performance.

The author-topic model has lower perplexity early on (for small values of N(train)d ) since it uses knowledge of the author to provide a better prior for the content of the document. However, as N(train)d increases we see a cross-over point where the more flexible topic model adapts better to the content of this particular document.

For larger numbers of topics, this crossover occurs for smaller values of N(train)d , since the topics pick out more specific areas of the subject domain.

One can see that making use of the authorship information significantly improves the predictive log-likelihood: the model has accurate expectations about the content of documents by particular authors.

Such a task requires computing the similarity between authors. To illustrate how the model could be used in this respect, we defined the distance between authors i and j as the symmetric KL divergence between the topics distribution conditioned on each of the authors:

The topic distributions for different authors can also be used to assess the extent to which authors tend to address a single topic in their work, or cover multiple topics. We calculated the entropy of each author's distribution over topics on the NIPS data, for different numbers of topics.

When compared to the LDA topic model, the author-topic model was shown to have more focused priors when relatively little is known about a new document, but the LDA model can better adapt its distribution over topics to the content of individual documents as more words are observed.