顯示具有 community detection 標籤的文章。 顯示所有文章
顯示具有 community detection 標籤的文章。 顯示所有文章

2016年7月7日 星期四

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

Jeong, Y. K. & Song, M. (2016). Applying content-based similarity measure to author co-citation analysis. In Proceedings of iConference 2016.

本研究利用引用文獻出現文句內容的相似性來測量作者的主題相關性(topical relatedness)。傳統的作者共被引分析(Author co-citation Analysis, ACA)做法是利用參考文獻裡被引用作者的共被引頻率(White and Griffith, 1981),然後利用Pearson相關係數 (Pearson correlation coefficient)或是 Salton提出的餘弦相似性測量作者的相似性,在書目計量學研究裡已經廣泛運用於確認與追蹤學科的知識結構(the intellectual structure of an academic discipline) (He & Hui, 2002)。然而這種做法並未考慮引用的內容,Jeong, Song, & Ding, (2014)與 Zhao & Strotmann (2014)則利用全文裡提到的作者並將有關的內容加入ACA的計算。

本研究認為累積被引作者出現的文句能夠代表作者的研究領域,因此利用JASIST的全文資料,剖析HTML,取出論文的後設資料(題名、作者姓名、出版年、DOI與摘要)、引用資訊(引用文句與參考文獻索引)以及參考資訊(作者姓名、出版年、題名與期刊)。在這個研究裡,共使用2003年1月到2015年6月的1910篇論文,合計77,408筆參考文獻。將引用文句與一般文句分開,連結文句內的參考文獻索引與參考資訊,選取100位最多被引用的作者,進行傳統的ACA以及本研究提出的新方法。本研究的新方法利用Mikolov et al., (2013)提出的Word2Vec 模型 (Word2Vec models),根據參考文獻出現的引用文句,找出作者間的相似性。Word2Vec 模型以大量的文本為基礎,利用類神經網路方法( neural network approaches),找出詞語之間的語意關係,將每一個出現於文句的詞語轉換成向量,使得這些向量之間的相似性能夠保持詞語在語意上的關係。本研究將被引用的作者姓名視為是引用文句中出現的詞語,測量作者間在研究主題的相似性與合作關係。

表2是傳統的ACA方法與本研究的方法分別找出的10組最相似的作者,本研究的方法找出10組最相似的作者中有一半是具有合作關係的作者。

另外,將兩種方法產生的作者關係分別繪製成網路圖,節點代表作者,利用PageRank決定的節點大小,節點的遠近由作者間的相似性決定,並且以Blondel, Guillaume, Lambiotte, & Lefebvre (2008)提出的社群偵測(community detection)方法進行分群。圖三與圖四分別是傳統ACA與本研究提出方法的結果。

圖三上可明顯地看到所有的作者分為兩群,依據社群偵測的分群結果,左邊的作者可再分為兩群:最左邊紅色的一群為研究資訊尋求行為(information seeking behavior)的作者,紫色的一群則與資訊檢索(information retrieval)研究有關,右邊綠色的一群則是研究書目計量學(bibliometrics)的作者。介於左右兩大群體的作者分別有兩位:Borgman和Salton。這兩位都是資訊科學領域傳統上會經常引用的作者。



在以Word2Vec方法產生的作者網路上,與資訊檢索有關的作者群組位於左方,包括上方的資訊尋求行為以及下方的文件檢索(document retrieval)兩個群組,書目計量學在圖四上則分為兩個有關的群組,一個主要包含作者分析(author analysis),另一則是期刊引用分析(journal citation analysis)與評鑑指標(evaluation indicator)。與圖三不同的是,圖四上的群組彼此間都有連結,並且圖形上更具體地呈現次學科(sub-disciplines)以及重要的作者。

Unlike other ACA studies, we used citing sentences to reflect topical relatedness of authors.

In  our  research,  we extended  traditional  approaches by  adopting Word2Vec, one  of  deep learning methods, to measure author similarity.

We also conducted in-depth network analysis of author maps.

The results of Word2Vec-based author map revealed more specific sub-disciplines and the important authors in perspective of topical influence than traditional approach does.

Author co-citation Analysis (ACA), which was introduced by White and Griffith (1981), has been widely used in bibliometrics researches to identify and trace the intellectual structure of an academic discipline (He & Hui, 2002). In ACA, traditional approaches relied on the co-citation frequency of cited authors in the reference section.

Thus, one of the main topics in ACA was methodological discussion of what kind of measure is appropriate and relevant for calculation of author similarities (Leydesdorff, 2005; van Eck & Waltman, 2007). Existing approaches based on co-citation frequencies such as Pearson correlation coefficient and Salton’s cosine similarity, however, do not capture the citation content.

Thus, some recent researches used the full-text to obtain the topical relatedness between the cited authors (Jeong, Song, & Ding, 2014; Zhao & Strotmann, 2014). They analyzed the authors mentioned in the full-text and incorporated contents related with cited authors into ACA.

In that sense, cumulated citing sentences of cited authors are able to well represent the cited researches and cited authors’ research areas. In addition, these citing sentences are particularly useful for summarization of a research document.

Figure 1 shows the overall system flow of our approach.


For content analysis, however, we collected full-text research articles of JASIST in HTML format. Through the HTML parsing process, we extracted the metadata (title, author name, year, DOI and abstract), citation information (citing sentence, and reference id), and reference information (author name, year, title, and journal).

To compare our method to traditional ACA, we computed author-pairs in both approaches. In Word2Vec-based method, the full-text data, first, are splitting into sentences. In second step, matching the citing sentences with reference id in reference section, we separated the citing sentences and other general sentences. Then, citing sentences are preprocessed in the following steps: tokenization, POS tagging, lemmatization of the tokenized sentence, and stop word removal.

From these data, we trained Word2Vec model for calculating author similarity and generated author-author similarity matrix. To compare the previous research, traditional author counting approach, we also construct co-citation matrix based on citation counts. Since we preprocessed full-text including all reference information, these matrices considered all cited authors.

To evaluation, we selected top 100 authors which are highly cited in both methodology, and conduct network analysis through visualizing author maps.

The data was gathered from 1,910 full-text articles in the JASIST digital library over 12 years (from January 2003 to June 2015). The 1,910 collected documents have 77,408 references. We extracted elements from the full-text article: 1) citing sentences from the body of the article, 2) the references information, and 3) all cited authors. Table 1 shows the basic statistics of collected data.

Word2Vec models, one of the neural network approaches, are able to carry semantic meanings and turns text into a numerical form that deep-learning nets can understand (Mikolov et al., 2013). Based on a large amount of plain text, Word2Vec trains relationships between words automatically.

Word2Vec spatially encoded a word meaning and the relationship between words, which was originally applied to word clustering or synonym detection (Wolf et al., 2014). We applied Word2Vec into author similarity measure regarding cited author names as a word in plain text.

Since authors’ oeuvre was represented as the citing sentences in research articles, the Word2Vec-based method could consider various topics of the author.

In the proposed approach, however, the author names are also trained as words in a same citing sentence. Therefore, the similarity between two authors in the Word2Vec-based method reflects both topical relatedness and collaborations.

Table 2 shows top 10 pairs by the traditional ACA method (Pearson correlation based similarity) and the Word2Vec based approach respectively. About the half of pairs resulted from the Word2Vec approach are the co-author relationship.

This results imply that the proposed approach enables to detect wider range of author pairs in perspective of topical relatedness and grasp more diverse research fields of information science.

To examine whether there are structural differences in two measures of author similarity, we constructed two author networks with top 100 authors. For network visualization, we used PageRank (Brin & Page, 1998) to determine the node size and also adopted the modularity algorithm (Blondel, Guillaume, Lambiotte, & Lefebvre, 2008) for the community detection.

Figure 3 illustrates roughly two parts that consist of information retrieval and bibliometrics, two major research areas in JASIST. The author group of information retrieval (purple) along with information seeking behavior (red) is located at the left side, and the author group related with bibliometrics is located at the right side.

There are only two authors located between two groups (Borgman and Salton), who are traditionally cited authors in the information science field. Borgman studied various topics including information retrieval and scholarly communication and wrote the important books that had won the best information science book from ASIST. Salton’s works also received a lot of citations for a long time in the field of information science.


The author group related with information retrieval in the left side of the network is split into information seeking behavior (blue) located in the upper side of the network and document retrieval (yellow) located at the bottom side of the network. The group related to bibliometrics is also separated into two parts: (1) a cluster (green) including author analysis and (2) journal citation analysis and evaluation indicator (red).

Unlike Figure 3, the communities in the network are connected to each other. Brin is connected with both document retrieval and citation analysis communities. This may be attributed to the fact that the PageRank, developed by Brin and Page (1998), is used in information retrieval and also studied in network analysis to compute node centrality.

In bibliometrics, PageRank is adopted as one of the centralities in citation networks (Ding, Yan, Frazho, & Caverlee, 2009). Ingwesen, who is located between information retrieval and bibliometrics, studied information retrieval in earlier works, he extended the research area to network analysis such as webometrics.

It implies that the authors linked by citation are topically grouped in the Word2Vec-based author network.

2015年4月9日 星期四

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

Rafols, I., & Leydesdorff, L. (2009). Content‐based and algorithmic classifications of journals: Perspectives on the dynamics of scientific communication and indexer effects. Journal of the American Society for Information Science and Technology, 60(9), 1823-1835.

本研究比較兩種以內容為基礎的期刊分類以及兩種以演算法為基礎的期刊分類。兩種以內容為基礎的期刊分類分別是ISI的主題分類(Subject Categories)以及Glänzel and Schubert (2003)的領域/次領域分類(field/subfield classification)SOOI,兩種以演算法為基礎的期刊分類則分別是Blondel et al. (2008)提出的展開式(unfolding)社群偵測(community detection)法以及Rosvall, and Bergstrom (2008)的隨機漫步(random walk)矩陣分解(matrix decomposition)法。若是利用以內容為基礎的分類,期刊可以同時指定多個類別;以演算法為基礎的期刊分類則可以使類別內的引用(within-category citation)對類別內的引用(between-category citation)的比率最大化,也就是將期刊彼此之間的引用資料排列成矩陣,經過適當的行列排列後,使得主要對角線(principal diagonal)附近的數值較大,而其他地方則接近0。

各種分類的相關統計資料如表1所示:


由於以內容為基礎的分類方法具有多重分類特性以及以演算法為基礎的分類方法以矩陣分解為目的,從表1上可以觀察到兩種現象:1) 在類別內期刊數的中位數方面,可以看到以內容為基礎的兩種期刊分類方法較以演算法為基礎的期刊分類方法來得多,可配合圖1每個類別期刊數的分佈在0.50上所呈現的情形。另外,圖1也可發現四種分類方法都是對數常態分布(log normal distribution),也就是在這四種分類方法中,相對少數的類別擁有大量的期刊,然而許多類別卻只有少量期刊。並且以演算法為基礎的分類方法比以內容為基礎的分類方法更偏斜(more skewed),也較是上述的情況更嚴重。隨機漫步方法的前十個類別共有57%種期刊,展開方法則有50%,但ISI和SOOI則分別只有15%和31%。


2) 從引用的分布情形來看,兩種以內容為基礎的分類方法的引用次數總計比以演算法為基礎的分類方法多,但隨機漫步方法和展開方法有較多比率分布在類別內,但ISI和SOOI則是主要分布在類別之間。

接下來,以引用式樣(citation patterns)的餘弦相似性(cosine similarity),比較各種分類方法的類別彼此間的相似性。結果ISI和SOOI的中位數分別是0.020和0.066,比隨機漫步方法和展開方法的0.009和0.007高許多,其原因同樣是因為內容為基礎的方法有多重分類的特性,因此類別間的邊緣較模糊,而演算法為基礎的方法在類別間切割得較清楚。然後將各種分類方法的類別依照它們的相似性繪製成網路圖。四種方法繪製的網路圖大致上都可以看出包含兩大群,一個是生物醫學,另一個則是物理學與工程學,兩個大群體透過三個群體相連,包括化學、地理學-環境科學-生態學群體、以及電腦科學,社會科學群體在網路圖上有些分離,透過行為科學/神經科學和生物醫學相連,並且也透過電腦科學與數學和物理學/工程學相連。綜上所述,不同的科學地圖是相似的,但它們在群體內部類別的密度不同。

In this study, we test the results of two recently available algorithms for the decomposition of large matrices against two content-based classifications of journals: the ISI Subject Categories and the field/subfield classification of Glänzel and Schubert (2003).

The content-based schemes allow for the attribution of more than a single category to a journal, whereas the algorithms maximize the ratio of within-category citations over between-category citations in the aggregated category-category citation matrix.

At that time, Leydesdorff & Rafols (2009) were deeply involved in testing the ISI Subject Categories of these same journals in terms of their disciplinary organization. Using the JCR of the Science Citation Index (SCI), we found 14 major components using 172 subject categories, and 6,164 journals in 2006. Given our analytical objectives and the well-known differences in citation behaviour within the social sciences (Bensman,2008), we decided to set aside the study of the (220 − 175 = ) 45 subject categories in the social sciences for a future study.

Our findings using the SCI indicated that the ISI Subject Categories can be used for statistical mapping purposes at the global level despite being imprecise in terms of the detailed attribution of journals to the categories.

In this study, we compare the results of these two algorithms with the full set of 220 Subject Categories of the ISI. In addition to these three decompositions, a fourth classification system of journals was proposed by Glänzel and Schubert (2003) and increasingly used for evaluation purposes by the Steungroep Onderwijs and Onderzoek Indicatoren (SOOI) in Leuven, Belgium. These authors originally proposed 12 fields and 60 subfields for the SCI, and three fields and seven subfields for the Social Science Citation Index and the Arts and Humanities Citation Index. Later, one more subfield entitled “multidisciplinary sciences” was added.

Thus, because research topics are, on the one hand, thinly spread outside the core group and, on the other hand, the core groups are interwoven, one cannot expect that the aggregated journal-journal citation matrix matches one-to-one with substantive definitions of categories or that it can be decomposed in a single and unique way in relation to scientific specialties. The choice of an appropriate journal set can be considered as a local optimization problem (Leydesdorff, 2006).

Citation relations among journals are dense in discipline-specific clusters and are otherwise very sparse, to the extent of being virtually non-existent (Leydesdorff & Cozzens, 2003).

The grand matrix of aggregated journal-journal citations is so heavily structured that the mappings and analyses in terms of citation distributions have been amazingly robust despite differences in methodologies (e.g., Leydesdorff, 1987 and 2007; Tijssen, de Leeuw, & van Raan, 1987; Boyack, Klavans, & Börner, 2005; Moya-Anegón et al., 2007; Klavans & Boyack, 2009).

A decomposable matrix is a square matrix such that a rearrangement of rows and columns leaves a set of square sub-matrices on the principal diagonal and zeros everywhere else.

In the case of a nearly decomposable matrix, some zeros are replaced by relatively small nonzero numbers (Simon & Ando, 1961; Ando & Fisher, 1963). Near-decomposability is a general property of complex and evolving systems (Simon, 1973 and 2002).

The decomposition into nearly decomposable matrices has no analytical solution. However, algorithms can provide heuristic decompositions when there is no single unique correct answer.

Newman (2006a, 2006b) proposed using modularity for the decomposition of nearly decomposable matrices since modularity can be maximized as an objective function.

Blondel et al. (2008) used this function for relocating units iteratively in neighbouring clusters. Each decomposition can then be considered in terms of whether it increases the modularity.

Analogously, Rosvall, and Bergstrom (2008) maximized the probabilistic entropy between clusters by estimating the fraction of time during which every node is visited in a random walk (cf. Theil, 1972; Leydesdorff, 1991).

The data were harvested from the CD-Rom version of the JCR of the SCI and Social Science Citation Index 2006, and then combined. ... The resulting set of 7,611 journals and their citation relations otherwise precisely corresponds to the online version of the JCRs. This large data matrix of 7,611 times 7,611 citing and cited journals was stored conveniently as a Pajek (.net) file and used for further processing.

The 7,611 journals are attributed by the ISI with 11,856 subject classifiers. This is 1.56 (±0.76) classifiers per journal. The ISI staff assign the 220 ISI Subject Categories on the basis of a number of criteria including the journal's title and its citation patterns (McVeigh, personal communication, March 9, 2006; Bensman & Leydesdorff, 2009).

According to the evaluation of Pudovkin and Garfield (2002), in many fields these categories are sufficient, but the authors added that “in many areas of research these ‘classifications’ are crude and do not permit the user to quickly learn which journals are most closely related” (p. 1113).

Leydesdorff and Rafols (2009) found that the ISI Subject Categories can be used for statistical purposes—the factor analysis for example can remove the noise—but not for the detailed evaluation. In the case of interdisciplinary fields, problems of imprecise or potentially erroneous classifications can be expected.

For the purpose of developing a new classification scheme of scientific journals contained in the SCIs, Glänzel and Schubert (2003) used three successive steps for their attribution. The authors iteratively distinguished sets cognitively on the basis of expert judgements, pragmatically to retain multiple assignments within reasonable limits, and scientometrically using unambiguous core journals for the classification. The scheme of 15 fields and 68 subfields is used extensively for research evaluations by the Steunpunt Onderwijs and Onderzoek Indicatoren (SOOI), a research unit at the Catholic University in Leuven, Belgium, headed by Glänzel.

The SOOI categories cover 8,985 journals. Using the full titles of the journals, 7,485 could be matched with the 7,611 journals under study in the JCR data for 2006 (which is 98.3%). These journals are attributed 10,840 classifiers at the subfield level. This is 1.45 (±0.66) categories per journal. One category (“Philosophy and Religion”) is missing because the Arts & Humanities Citation Index is not included in our data. Thus, we pursued the analysis with the 67 SOOI categories.

Using Rosvall and Bergstrom's (2008) algorithm with 2006 data, we obtained findings similar to those of these authors on August 11, 2008. Like the original authors using 6,128 journals in 2004, we found 88 clusters using 7,611 journals in 2006.

Lambiotte, one of the coauthors of Blondel et al. (2008), was so kind as to input the data into the unfolding algorithm and found the following results: 114 communities with a modularity value of 0.527708 and 14 communities with a modularity value of 0.60345. We use the 114 communities for the purposes of this comparison. These categories refer to 7,607 (= 7611 − 4) journals because four of the journals in the file were isolates.

The number of journals per category is log-normally distributed in each of the four classifications. In other words, they all have a relatively small number of categories with a large number of journals and many categories with only a few journals. However, as shown in Figure 1, the classifications based on the random walk and unfolding algorithms are more skewed than the content-based classifications.



Whereas the top-10 categories on the basis of a random walk comprise 57% of the journals (50% for unfolding), they cover only 15% in the ISI decomposition and 31% for the SOOI classification. In the case of skewed distributions, the characteristic number of journals per category can best be expressed by the median: the median is below 30 in the random walk or unfolding classifications, compared with 42 journals for the ISI classification and 141 for the SOOI classification (Table 1).


As presented in the last rows of Table 1, the total numbers of citations in the aggregated matrices based on the ISI or SOOI classifications are much higher because the same citation can be attributed to two or three categories. Thus, whereas random walk and unfolding lead to matrices with most citations within categories (on the diagonal), matrices based on ISI and SOOI classifications lead to matrices with most citations between categories (off-diagonal).

Finally, to measure how similar the categories in the four decompositions are to each other, we computed the cosine similarities in the citation patterns between each pair of citing categories in the four aggregated category-category matrices (Salton & McGill, 1983; Ahlgren, Jarneving, & Rousseau, 2003).

We find again that all the distributions are highly skewed and that the random walk and unfolding algorithms exhibit a much lower median similarity value among categories. The lower medians indicate that the algorithmic decompositions produce a much “cleaner” cut between categories than the content-based classifications.
In conclusion, the analysis of the statistical properties of the different classifications teaches us that the random walk and the unfolding algorithms produce much more skewed distributions in terms of the number of journals per category, but these constructs are more specific than the content-based classification of the ISI and SOOI. The content-based sets are less divided because the boundaries among them are blurred by the multiple assignments.

In summary, although the correspondences among the main categories are sometimes as low as 50% of the journals, most of the mismatched journals appear to fall in areas within the close vicinity of the main categories. In other words, it seems that the various decompositions are roughly consistent but imprecise.

Maps of science for each decomposition were generated from the aggregated category-category citation matrices using the cosine as similarity measure.

The similarity matrices were visualized with Pajek (Batagelj & Mrvar, 1998) using Kamada and Kawai's (1989) algorithm.

The threshold value of similarity for edge visualization is pragmatically set at cosine > 0.01 for the algorithmic decompositions and cosine > 0.2 for the content-based decompositions to enhance the readability of the maps without affecting the representation of the structures in the data.

For the ISI decomposition, the 220 categories (Figure 3) were clustered into 18 macro-categories (Figure 4) obtained from the factor analysis (cf. Leydesdorff and Rafols, 2009).


The map of the SOOI classification was constructed with all is 67 subfields (Figure 5).


Taking advantage of the concentration of journals in a few categories, in the case of random walk and unfolding only the top 30 and 35 categories were used, respectively.


Indeed, the four maps correspond in displaying two main poles: a very large pole in the biomedical sciences and a second pole in the physical sciences and engineering. These two poles are connected via three bridging areas: chemistry, a geosciences-environment-ecology group, and the computer sciences. The social sciences are somewhat detached, linked via the behavioral sciences/neuroscience to the biomedical pole, and via the computer sciences and mathematics to the physics/engineering pole.

As noted above, although categories of different decompositions do not always match with one another, most “misplaced” journals are assigned into closely neighbouring categories. Therefore, the error in terms of categories is not large and is also unsystematic. The noise-to-signal ratio becomes much smaller when aggregated over the relations among categories.

As a second important observation that can be made on the basis of these maps, we wish to point to the differences in category density between the content-based and the algorithm-based maps.

In summary, we were surprised to find that the different science maps are similar except that they differ in the density of categories within groups.

The content-based classifications achieve a more balanced coverage of the disciplines at the expense of distinguishing categories that may be highly similar in terms of journals.

The first finding is that the algorithmic decompositions have very skewed and clean-cut distributions, with large clusters in a few scientific areas, whereas indexers maintain more even and overlapping distributions in the content-based classifications.

Second, the different classifications show a limited degree of agreement in terms of matching categories. In spite of this lack of agreement, however, the science maps obtained are surprisingly similar; this robustness is due to the fact that although categories do not match precisely, their relative positions in the network among the other categories is based on distributions that match sufficiently to produce corresponding maps at the aggregated level.

2014年6月21日 星期六

Jensen, P., & Lutkouskaya, K. (2014). The many dimensions of laboratories’ interdisciplinarity. Scientometrics, 98(1), 619-631.

Jensen, P., & Lutkouskaya, K. (2014). The many dimensions of laboratories’ interdisciplinarity. Scientometrics, 98(1), 619-631.

Scientometrics

本研究提出六種指標來測量研究機構的跨學科性。最廣義的來說,跨學科性可以視為不同學科某種程度的整合 (Weingart and Stehr 2000; Porter and Rafols 2009; Marcovich and Shinn 2011; Wagner et al. 2011; Rafols et al. 2012),為了將這個想法轉換為量化的指標,本研究認為需要考慮三個問題:
1. 如何定義一個學科
2. 在什麼層次達到整合
3. 學科連結需要到達什麼程度

在學科的定義上,本研究提出三種方式:一、因為是分析CNRS的實驗室,自然可採用CNRS的學科組織(disciplinary organization),包括10個研究所(institutes)以及進一步細分成的40個組(sections);二、如同其他先前的研究,使用WoS(Web of Science)的224種期刊主題分類(Journal Subject Categories, JSCs);三、將文件根據共同的參考文獻,以叢集演算法(clustering algorithms)由下往上地(bottom-up)歸類成認知叢集(cognitive clusters)。在整合的層次,則探討實驗室與論文兩個層級。

以實驗室的跨領域程度來說,較簡單的方式可以定義為:
此處的pi是實驗室的論文在期刊主題分類JSC i上的比例。

除了上述的定義之外,本研究還使用的Stirling’s (2007)方法來表現多樣性的三個不同面向:不同類別的數量(variety)、在各類別上的分布均勻程度(balance)、以及表現類別間的差異(disparity) (Porter and Rafols 2009):
此處的是主題分類JSC i 和 JSC j的相似性,並且此一相似性以cosine測量主題分類間的引用情形得到。(Porter and Rafols 2009).

為了進一步了解實驗室的跨領域多樣性是否在單一論文的認知層次達成,如同上述的情形,計算單一論文的跨領域多樣性時,可以利用下面的方式:
此處的pai是此論文引用的參考文獻在期刊主題分類JSC i上的比例。進一步用實驗室發表的論文考慮實驗室的跨領域程度時,可以將所有論文的跨領域多樣性加以平均,如

此處的 #pap 是實驗室發表的論文數量。

此外,另兩種指標分別是主流引用外的主題分類比例以及不同機構的人員合作占論文全體比率,分別如下所示:


最後一種指標,先以書目耦合(bibliographic coupling) (Kessler 1963)產生論文之間的關連,計算方式如下:
此處的where #common_refsij 是論文 i 和 j 共同引用的參考文獻數量, #refsi 和 #refsj 分別是論文 i 和 j 包含的參考文獻數量。接下來以書目耦合關連建立論文網路,希望在網路上引用文獻相似的論文會聚集形成叢集。因此,接下來Blondel et al. (2008)的演算法,劃分網路成論文的叢集。整個方法可參見Grauwin and Jensen (2011),結果共劃分成250個叢集。然後以下面的方式計算實驗室在認知叢集上的多樣性
此處的 p_i 和 p_j 分別是實驗室的論文屬於叢集 i 和 j 的比例。

以六種指標計算每一個實驗室的跨學科多樣性後,接下來以主成分分析(Principal Component Analysis, PCA)進行分析,四個主要的成分分別是
1) 實驗室在各種多樣性指標的綜合表現
2) 實驗室連結的學科的認知距離(cognitive distance)
3) 實驗室在實驗室層級或論文層級具有跨學科性
4) 論文發表的期刊具有跨學科的主題分類或是與其他不同機構的實驗室合作。

Interdisciplinarity is as trendy as it is difficult to define. Instead of trying to capture a multidimensional object with a single indicator, we propose six indicators, combining three different operationalizations of a discipline, two levels (article or laboratory) of integration of these disciplines and two measures of interdisciplinary diversity.

Interdisciplinarity means, at the most generic level, some degree of integration of different disciplines (Weingart and Stehr 2000; Porter and Rafols 2009; Marcovich and Shinn 2011; Wagner et al. 2011; Rafols et al. 2012).

To transform this idea into quantitative indicators, we need to answer three questions:
1. How to define a discipline?
2. At what level the integration is achieved?
3. What is the degree of disciplinary linkage achieved?

There are several ways to define a discipline from a scientometrics’ point of view. Since we are dealing with CNRS labs, the most natural would seem to use the disciplinary organization of CNRS in 10 ‘‘institutes’’ and 40 subdisciplinary ‘‘sections’’. A convenient alternative is to use the 224 Journal Subject Categories (JSCs) used by Web of Science (WoS). Finally, instead of using institutionally predefined divisions of science, one could use a more bottom-up definition of ‘‘cognitive clusters’’. To obtain these clusters, we use the roughly 300,000 French articles published between 2007 and 2010 and group them into ‘‘cognitive clusters’’ using clustering algorithms based on shared references.

In this paper, we will use three definitions of ‘‘discipline’’ and two integration levels (laboratory and article) to calculate six partial interdisciplinary indicators.

We adopt Stirling’s (2007) approach to capture the different facets of diversity : ‘variety’, ‘balance’ and ‘disparity’.

‘Variety’ characterizes the number of different categories, ‘balance’ characterizes the evenness of the distribution over these categories and ‘disparity’ characterizes the difference among the categories, usually based on some distance.

A simple indicator of the spread of the disciplines where a laboratory publishes is given by:
where pi is the proportion of articles of the laboratory in JSCi.

As we would like to include the idea of ‘‘distance’’ between disciplines, we calculate the diversity indicator (Stirling 2007; Porter and Rafols 2009) which combines both the spread of the disciplines through the pi and the distance between them.
where sij is the cosine measure of similarity between JSCs i and j. Practically, sij is measured through the citations from publications in JSCsi to publications in JSC j (Porter and Rafols 2009).

To further characterize a lab’s interdisciplinarity, it is useful to introduce an indicator of the interdisciplinarity of single articles, to test whether interdisciplinarity is achieved at this cognitive level.

Specifically, the interdisciplinary diversity of a single article is calculated as:
where pai is the proportion of articles’ references in JSCi.

To quantify the interdisciplinarity of the papers published by a lab, we aggregate the articles’ diversity indicator art_div_corr at the laboratory level by averaging over all the articles published by that laboratory:

where #pap is the number of articles of the lab for which at least one reference was identified.

Then, we choose a threshold to define the most common JSCs for each institute. ... We therefore choose a threshold value of 90 %. ... Then, for each laboratory, we count the percentage of articles outside this 90 % list and normalize by the expected value, i.e. the average value 0.1.


whereare the frequencies of the JSCs that do not belong to the Institute’s JSC main list.

Interdisciplinary collaborations can also be detected by copublications between scientists belonging to different CNRS Institutes. We compute a fifth indicator by calculating the proportion of a lab’s publications that involve authors from other Institutes

where the sum counts the number of articles of the lab involving at least two institutes and
#articles is the total number of articles published by the laboratory.

To build these ‘‘cognitive disciplines’’, we use bibliographic coupling (BC) (Kessler 1963) between the 300,000 papers published by French laboratories in the period 2007–2010 and compiled by the WoS.
where #common_refsij is the number of common references for articles i and j, and #refsi,
#refsj are the numbers of references of articles i and j, respectively.

In comparison to a co-citation link (which is the usual measure of articles’ similarity), BC offers two advantages: it allows to map recent papers (which have not yet been cited) and it deals with all published papers (whether cited or not).

This reinforcement facilitates the partition of the network into meaningful groups of cohesive articles, or clusters. A widely used criterion to measure the quality of a partition is the modularity function (Fortunato and Barthe´lemy 2007), which is roughly is the number of edges ‘inside clusters’ (as opposed to ‘between clusters’), minus the expected number of such edges if the partition were randomly produced. We compute the graph partition using the efficient heuristic algorithm presented in (Blondel et al. 2008). The whole method is described in (Grauwin and Jensen 2011).

Applying this algorithm yields in a partition of French papers into roughly 250 clusters containing more than 100 papers each.
where p_i is and p_j are the proportions of the labs’ papers belonging to clusters i and j respectively.

On average, articles refer to papers from almost 10 different disciplines (9.8 JSC). .... However, when considering those JSC that are used in more than 10 % of the reference list, this average drops to 2.7. This means that, on average, an article spreads its references on 3 main JSCs and 7 additional which benefit from roughly a single reference.

An average laboratory publishes in journals belonging to 34 different JSCs ...

PCA1: combined interdisciplinarity The main axis represents a combination of the various interdisciplinarity indicators.

PCA2: short or long cognitive distance This axis distinguishes those labs that connect distant or nearby disciplines.

PCA3: article or laboratory interdisciplinarity This axis distinguishes labs that achieve interdisciplinarity either at the laboratory or article level.

PCA4: diversity of publications’ JSCs or diversity of collaborations This axis distinguishes labs that publish in journals belonging to different JSCs (high lab_jsc_bal) from labs that co-publish with labs from different CNRS Institutes (high lab_inst_cop_bal).

We have computed the six indicators for the 680 laboratories which have published more than 50 papers over 2007–2010. To allow comparisons and statistical analysis, since the absolute values have no intrinsic meaning, we have scaled all the values to achieve an average value of 0 and a variance of 1. We then carried out a principal component analysis of the (680 9 6) matrix using the free software R (www.r-project.org/). More precisely, we used prcomp from the ‘stats’ package, without any axes rotation.

First, let us note that using the first four PCA axes gives an overall view about the interdisciplinarity practices of each lab. This view has been compared to expert knowledge, namely scientists working in those labs or scientific advisors from CNRS. This comparison, carried out for about 20 different labs from all the disciplines, suggests that these indicators characterize interdisciplinarity
in a meaningful way.

A major drawback of our method is that we cannot distinguish real interdisciplinary collaborations, giving rise to new concepts or to a coherent new scientific field, from simple pluridisciplinary practices that merely juxtapose different disciplines, as when historians use characterizing tools from physics. It seems difficult to learn much about the cognitive dimensions of interdisciplinarity from an automatic analysis of metadata of the papers.

2013年12月7日 星期六

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

Li, D., He, B., Ding, Y., Tang, J., Sugimoto, C., Qin, Z., ... & Dong, T. (2010, October). Community-based topic modeling for social tagging. In Proceedings of the 19th ACM international conference on Information and knowledge management (pp. 1565-1568). ACM.

本研究提出一個TTR-LDA-社群模型,這個模型以推論機制(inference mechanism)結合LDA(Latent Dirichlet Allocation)模型和Girvan-Newman社群偵測(community detection)演算法提供在網路資料上偵測社群並對這些社群進行主題探勘(topic mining)的功能,並且進而了解在社群上的主題隨時間推移的變化,處理的架構如下圖所示

本研究利用Delicious社會標籤系統(social tagging system)上從2005到2008年的資料進行研究。在社群偵測部分,首先建立網絡:根據使用者標籤的資源數量,選取前50000位標籤資源最多的使用者;然後對他們標籤的網頁進行統計,從其中選取10000個被最多使用者標籤的網頁。接著在上述的10000個網頁中,如果有這50000位使用者之間有任何兩位曾經標籤過相同的網頁,便在這兩位使用者之間產生一個連結。以50000個使用者為節點,同時以他們之間的連結為連結線,便可以建立一個共同書籤網絡(co-bookmark network)。並且為了研究網絡上社群結構的變化,並將整個期間的資料分為三個時段:分別為2005-2006、2007與2008年,相關的統計數據如下表:

本研究利用Girvan-Newman演算法找出標籤者(Tagger)社群,使得標籤者與同一社群內的其他標籤者比社群外的標籤者有較強的關係。這個演算法重複移去當時網絡上中介性(betweenness)最大的連結線,產生各種可能的網路劃分(network partition),測量每一種劃分下的群組性(modularity),也就是實際上社群內的成員彼此間的連結線數量與相同連結度的情況但隨機產生連結線的數量的差,群組性最大的劃分便是輸出結果。

另一方面,本研究利用TTR-LDA模型找出每個標籤者的主題分布以及主題內具有代表性的標籤,TTR-LDA模型修改自ACT(author-conference-topic)模式[11][12],是一個由標籤者做為第一層、標籤與資源為第三層、主題則為第二層,所構成的三層貝氏模型(three-layer Bayesian model)。

整合Girvan-Newman演算法和TTR-LDA模型的方法是以社群為單位,將社群內所有標籤者的主題分布進行平均做為該社群的主題分布,根據主題分布,選出機率值較大的主題做為社群的代表。比較不同時段社群共同的代表標籤衡量它們的相似性。

社群偵測的結果發現前五個最大的社群在四年裡占了絕大多數的比率,而且這個比率逐年增加。主題探勘的部分則測量TTR-LDA模型在不同主題數量上的複雜度(perplexity),發現150個主題時有最低的複雜度。因此,以下的研究便針對150個主題在前五個最大社群上的分布進行探討,計算它們的傳導性(conductance)與模組性。就社群模組性而言,最近一個時段(2008年)的結果比前三個時段還要高,其原因可能是因為經過一段時間後,社群的結構逐漸成熟,因此後期比前期更能產生較佳的社群。另外,將最後一個時段再細分為四個較小的時段則發現,較小時段的社群模組性比整年的結果來得高,本研究認為造成這種現象的原因可能是由於在不同的時段,大部分標籤者的書籤行為集中在不同的領域;當那些時段合併起來的時候,會展現標籤者在多個領域的興趣,使得社群內的叢集(clustering)特性較弱。

結果並可以發現前二十個主題都出現在不同時段的前五個最大社群裡,每個社群至少包括一個前十名的主題。此外,比較LDA、TTR-LDA和TTR-LDA-社群等三種模型在資源和標籤上的預測力,在回收率(recall rate)、精確率(precision)和F1指標上以TTR-LDA-社群為最佳。

In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism.

The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008.

Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time.

From a research perspective, these real-world networks display unique properties from the classical random graph model [3] in that most real word networks exhibit three common properties: the small-world property, power-law degree distribution and a high clustering coefficient or transitivity (indicating community structure) [7][8][9].

Thus, an important task in network analysis is to detect communities and explore their features, which can improve community-supporting services at the community-level in the context of a social tagging system.

Many studies in various disciplines have been devoted to community detection; however, few of them have systematically and quantitatively studied the profiles of those detected communities.

In this paper, we propose a TTR-LDA-Community model, which is an inferential combination of an extended LDA model and a betweenness-based community detection algorithm. It provides rich, systematic, and quantitative information about the profiles of detected communities.

In the context of social tagging systems, where multiple users are annotating resources, the resulting topics reflect a shared view of the document; and the tags of the topics reflect a common vocabulary.

Girvan and Newman extended the betweenness measure to edges and designed a clustering algorithm which gradually removes the edges with the highest betweenness value [4]. This algorithm has been improved through modularity; and the complexity is reduced from O(m2n) to O(mdlogn) where d is the depth of the dendrogram of the community structure [2].

Many studies provide various models and algorithms for topic mining and community detection; yet, few of them have integrated those models and algorithms, performed topic mining for detected communities, and analyzed how those identified topics change among communities over time.

The activity of social tagging consists of three major components: tag, tagger and resource. The experimental dataset contains all the triples of these three components and the time and date of their creation on Delicious from 2005 to 2008.

In data processing, all taggers were ranked by the number of resources they have bookmarked and the top 50,000 taggers were selected as the sample of taggers.

These taggers bookmarked a total of 354,522 web pages, which were sorted by the number of taggers who bookmarked them. The top 10,000 resources were selected as the sample of web pages, associated with which a dominant majority of tagging activities occurred.

Thus a co-bookmark network was built in which a connection between two users (within the sample of 50,000 taggers) is created if they bookmarked the same resources (within the sample of 10,000 web pages).

In addition, in order to observe the evolution of structure and motif of communities, the time span (2005-2008) was divided into three slices.


The model is illustrated in Figure 1. TTR-LDA is developed based on ACT model [11][12]. It is a three-layer Bayesian model with taggers tap in each post p as the first layer, tags t, and resource r as third layer and all the topics denoted as latent variable z as the middle layer.

The inference mechanism is used to infer the topic distribution over detected communities.

Each community includes a set of taggers, who have a stronger relationship with other taggers within the community than the taggers outside.

Based on the taggers’ information model, the probability distribution of each tagger over a set of topics is obtained by using the TTR-LDA model while the community structure of taggers is revealed by the community detection algorithm. The two sets of results are further integrated through an inference mechanism.




Results show that the number of users of the top five communities occupies a major proportion in the four years (2005-2008) and the proportion is increasing over time.

Perplexity is used to identify the number of topics [10], which arrives at the lowest point when the number of topics is 150. The interest model of each tagger in the top five largest communities is then built based on their topic distributions.

By using users’ interest models and the inference mechanism, a topic distribution of the largest community can be created (Figure 3). We can find that the topic distributions in a community are diverse because users’ relationships in that community are mainly based on their co-bookmark activities not the similarity of their interest model.

In order to observe the dynamic features of communities, we design an experiment as follows:
1) denote the five largest communities from each time slice in 2008 as community_i_t where t means the tth time slice in 2008 and i means the ith largest community in tth time slice;
2) compute the topic distribution for the five communities, which is stored as model_t_i_Topic(j), the
probability of jth topic in ith largest community in the tth time slice;
3) obtain the probability distribution of tags that are collected from all the posts generated during the specific time slice; the probability of one tag occurring in a topic shows the level of representativeness of the tag for that topic;
4) sort all the tags according to their probability value in each topic and select the 20 top ranked tags to represent the content of the topics; select the top 5 ranked topics to represent the theme of each community;
5) analyze the similarity between different communities from different time slice through computing how many tags are shared by the two different communities. More specifically, we compare current time slice with its previous time slice, for example, we compare community_i_t with community_j_t-1 (j=1, 2…5).

The size of communities along evolutionary lines fluctuates over time. For example, the size of the community about social networks in the 3rd time slice (community_1_3) is much larger (4,377) than that (521) in the 4th time slice (community_5_4).

Conductance (from multi-criterion scores) and modularity (from single criterion scores) are used to evaluate the quality of communities detected by the TTR-LDA-Community model [6].

The smaller the value of conductance is, the higher the granularity of a community is. Network community profile (NCP) is used to compute and display the value of conductance for communities [5].

Whiskers networks and rewired networks are adopted as two comparative aspects. Whiskers is defined as the maximal sub graphs that can be detached from the rest of the network by removing a single edge; and a rewired network is a random network that has the same nodes and the same degree distribution as the original network [5].

The conductance of communities of the rewired original network (blue line in the left figure), rewired random network (red dashed line in the left figure), the original whiskers network (blue line in the right figure), and the random whiskers network (red dashed line in the right figure) are calculated and shown in Figure 4.

In Figure 4, compared with the rewired network (left) and the rewired whiskers (right), 1) the original network displays a higher granularity of communities (a lower conductance value); 2) the value of conductance as the function of the size of communities in the original network and the original whiskers present a “V” shape, showing properties of a true large social networks [5]; 3) the original whiskers has the best community granularity (the lowest conductance) between size 10-100; and 4) the best community granularity of rewired original network is around 1000.

The modularity of communities in the four time slices of 2008 is better than that in 2005-2007. This is probably due to the fact that community structure grows mature gradually over time, creating better communities in later years than in earlier years.

Meanwhile, modularity of communities in the short-term (four sub periods in 2008) is larger than the long-term (2008). It can be explained that in different time periods, most taggers’ bookmarking activities are focused on different domains, so in a certain short-term time period, communities may be quite different from each other. However, when those time periods are merged together, the taggers show different interests in many domains; so the clustering feature within the communities becomes weaker.

Results show that the most popular topics are about bandslash fiction, fan fiction, and supernatural fiction (the top 3 popular topics). Communities with similar theme are ranked 3rd, 4th, and 5th in size; and the web resources with similar topics are ranked 500-600 of the top 1000 ranked resources in number of taggers associated with them.

The top 20 ranked topics in 1000 most popular resources can be found in 5 largest communities in different time periods. For each community, there exists at least one topic that is ranked top 10 in 1000 most popular resources (Table 3).

Topic distributions for each community are obtained respectively from LDA, TTR-LDA model, and TTR-LDACommunity model based on co-bookmark network in a given period (Oct. 2008–Dec. 2008). One resource and five tags are recommended for each post according to the results of three models separately.

The TTR-LDA and TTR-LDA-Community model show significant improvement for recommendation of tags and resources for post in terms of precision, recall and F1-Measure. TTR-LDA and TTR-LDA-Community have slightly improved performance for “tags for post”, while TTR-LDA-Community outperforms TTR-LDA on “resource for post”.


2013年11月29日 星期五

Zhang, H., Qiu, B., Giles, C. L., Foley, H. C., & Yen, J. (2007, May). An LDA-based community structure discovery approach for large-scale social networks. In Intelligence and Security Informatics, 2007 IEEE (pp. 200-207). IEEE.

Zhang, H., Qiu, B., Giles, C. L., Foley, H. C., & Yen, J. (2007, May). An LDA-based community structure discovery approach for large-scale social networks. In Intelligence and Security Informatics, 2007 IEEE (pp. 200-207). IEEE.

本研究以LDA (latent Dirichlet allocation) 演算法[17]做為社群發現(community discovery)的方法,也就是找出社會網路上在群體內高度相連但在群體間相對稀疏的行動者(actors)社群。在這個方法裡,社群被視為是LDA模型裡的隱藏變數,由所有行動者依不同比例組合成的混合(mixtures)。本研究認為這個方法的優點是只需要利用網路的型態資訊(topological information)便可進行社群發現;此外,有別於[2, 3, 5]等先前的方法,這個方法能夠描述行動者屬於多個社群的現象,並且在每個社群上有不同重要性。

本研究將所有與某個行動者互動的行動者以及他們之間的社會互動情形定義為這個行動者的社會互動特徵(social interaction profile, sip),例如行動者vi的社會互動特徵定義為

SIW(vi, wij)是行動者vi與另一位行動者wij間的互動情形,mi是與vi互動的行動者數目。
在利用LDA模型進行社群發現時,將每個行動者的社會互動特徵視為是原先模型中的文件,這些行動者發生的互動則是模型中的詞語,值得一提的是在社會網路上行動者間的互動順序是可以交換的(exchangeable),因此LDA模型相當適合這種情形。下圖是這個模型的圖示,


對行動者vi而言,其社會互動特徵sipi的產生過程(generative process)如下
1) Sample mixture components ϕk ~ Dir(β) for k (belongs) [1, K]
2) Choose θi ~ Dir(α)
3) Choose Ni ~ Poisson(ξ) (note that Poisson assumption is not critical to this model)
4) For each of the Ni social interactions wij :
(a) Choose a community ιij ~ Multinomial(θi);
(b) Choose a social interaction wij ~ Multinomial(ϕιij)
也就是
1) 從Dir(β)中產生行動者的組合ϕk來代表K個社群中每一個社群;
2) 從Dir(α)中選擇一個社群的組合θi,做為所有與vi互動的行動者所可能屬於的社群的分布情形;
3) 從Poisson(ξ)中選擇社會互動的個數Ni ;
4) 產生每一個社會互動 wij時:
a) 從Multinomial(θi)中選擇一個社群ιij;
b) 從Multinomial(ϕιij)中選擇一個社會互動 wij。

根據這個模型,行動者vi的社會互動特徵sipi中的第j個社會互動元素wij為wm的機率是

此處θi是sipi 的混合比變數(mixing proportion variable),而ϕk是第k個社群成分分布(component distribution)的參數集合。

給定超參數(hyperparameter)α和β後,所有已知和隱藏變數的聯合機率(joint probability)為

由於要精確地推論LDA模型的變數一般而言是很困難的(intractable),因此經常以 變異期望值最大化(variational expectation maximization) [17], 期望值延遲 (expectation propagation) [24]和 Gibbs取樣 (Gibbs sampling) [25], [26], [22]等三種方法取得近似解。本研究採用的方法是Gibbs取樣,以Markov鏈 Monte Carlo模擬(Markov-chain Monte Carlo simulation)的方式,去逼近ϕk,w和θm,k等變數


本研究針對CiteSeer和NanoSCI兩個書目資料集中的作者合著網路(co-authorship networks)上最大的成分進行社群發現,其中CiteSeer部分共包含249866個節點,而NanoSCI則有203762個節點。本研究並且提出了01-SIP、012-SIP和 k-SIP等三種描述行動者間的社會互動情形,01-SIP是描述兩個作者有直接合著的社會互動情形,如果兩個作者曾經合著至少一篇論文,便將他們的互動情形設為1,否則便設為0;012-SIP則考慮兩個作者沒有直接合著但曾經分別和第三位作者合著的情形,如果兩個作者曾經合著至少一篇論文,便將他們的互動情形設為2,雖然這兩位作者不曾合著,但曾有分別和第三位作者合著,便將互動情形設為2,否則便設為0;k-SIP則描述兩位作者多次合著的情形,k是他們合著的次數。研究結果發現在複雜度(perplexity)與各產生社群的緊密度(compactness of communities)等指標,012-SIP的成效都比其他兩者好,從社群發現的結果可以發現許多社群裡的作者是屬於相同的研究機構或是彼此具有相似的研究興趣,因此會共同研究並且合作發表而成為社群。

This paper describes an LDA(latent Dirichlet Allocation)-based hierarchical Bayesian algorithm, namely SSN-LDA(Simple Social Network LDA). In SSN-LDA, communities are modeled as latent variables in the graphical model and defined as distributions over the social actor space. The advantage of SSN-LDA is that it only requires topological information as input.

This model is evaluated on two research collaborative networks: CiteSeer and NanoSCI. The experimental results demonstrate that this approach is promising for discovering community structures in large-scale networks.

An important task in these emerging networks is community discovery, which is to identify subsets of networks such that connections within each subset are dense and connections among different subsets are relatively sparse.

Unlike those previous community discovery studies, we design a hierarchical Bayesian network based approach, namely SSN-LDA (Simple Social Network-LDA) to discover probabilistic communities from social networks. ... In this model, communities are modeled as latent variables and are considered as distributions on the entire social actor space.

We also propose three different approaches to create social interaction profiles based on the social interaction information in the network.

Girvan et al extended this measure to edges and designed a clustering algorithm which gradually remove the edges with highest betweenness value [5]. ... However, a major problem with this approach is that the complexity of this approach is O(m2n), where m is the number of edges in the graph and n is the number of vertices in the network.

The graph partition problem can be formulated as the balanced minimum cut problem where the goal is to find an optimal graph partition so that the edge weight between the partitions is minimized while maintaining partitions of a minimal size. The NP-complete complexity of this approach [15] requires approximate solutions. Flake et al developed approximate algorithms to partition the network by solving s-t maximum flow techniques [2], [3]. The main idea behind maximum flow is to create clusters that have small inter-cluster cuts and relatively large intra-cluster cuts.

The major difference between SSN-LDA approach and the aforementioned approaches is that SSN-LDA is a mixture-model based probabilistic approach. Each community weighs in the contributions from every social actors and this property can be exploited in many potential applications that will be introduced in Section VI.

LDA model was first introduced by Blei for modeling the generative process of a document corpus [17].

Among these variants of LDA models, the approaches proposed in [20], [10] are both concerned about the authors of the documents in the corpus. In particular, Zhou et al introduced a community latent variable in their graphical model and applied it to discover community information embedded in document corpus. This approach can discover the underlying social network based on social interactions and topical similarity.

In their follow-up work[23], Zhou et al. investigated how research topics evolve over time and attempted to discover the most influential researchers involved in such transitions.

Each actor is characterized by its social interaction profile (SIP), which is defined as a set of neighbor(wij) and the corresponding weight(SIW(vi, wij)) pair.

where mi is the size of vi's social interaction profile.

Note that we consider the social interaction elements in this profile are exchangeable and therefore their order will not be concerned. It is this exchangeability that permits the application of LDA model [17].

Subsequently, we specify that a social network contains a set of communities ι (ι1, ι2, …, ιk) and each community in ι is defined as a distribution on the social actor space. In SSN-LDA, community assignments are modeled as a latent variable (ι) in the graphical model. The community proportion variable (θ) is regulated by a Dirichlet distribution with a known parameter α. Meanwhile, each social actor belongs to every community with different probabilities and therefore its social interaction profiles can be represented as random mixtures over latent communities variables.

The SSN-LDA model for social network analysis is illustrated in Fig. 1. Note that SSN-LDA resembles topic-based LDA model[17], with the social network being analogous to the corpus, the social interaction profiles being analogous to documents; and the occurrence of social interactions being analogous to words.

The distribution of topics in documents and the terms over topics are two multinomial distributions with two Dirichlet priors, whose hyperparameters are α and β respectively.

The dimensionality K of the Dirichlet distribution, which is also the number of community component distributions, is assumed to be known and fixed.

This generative process for an agent(wi)'s social interaction profile sipi in a social network is:
1) Sample mixture components ϕk ~ Dir(β) for k (belongs) [1, K]
2) Choose θi ~ Dir(α)
3) Choose Ni ~ Poisson(ξ) (note that Poisson assumption is not critical to this model)
4) For each of the Ni social interactions wij :
(a) Choose a community ιij ~ Multinomial(θi);
(b) Choose a social interaction wij ~ Multinomial(ϕιij)

According to the model, the probability that the jth social interaction element wij in the social actor wi's social interaction profile sipi instantiates a particular neighboring agent wm is:

where θi is the mixing proportion variable for sipi and ϕk is the parameter set for the kth community component distribution.

Given the hyperparameters α and β, the joint distribution of all known and hidden variables is:


Exact inference is generally intractable for LDA model. There have been three major approaches for solving this model approximately, including variational expectation maximization [17], expectation propagation [24], and Gibbs sampling[25], [26], [22].

Gibbs sampling is a special case of Markov-chain Monte Carlo (MCMC) simulation[27] where the dimension K of the distribution are sampled alternately one at a time, conditioned on the values of all other dimensions[22].

we apply the Gibbs sampling algorithm that has been introduced in [26], [22] to solve the SSN-LDA model and reduce the computation requirement.



Subsequently, the update equation for the hidden variable can be derived [22]:


where n(.)⌐i is the count that does not include the current assignment of ιi and recall that sip is the variable for social interaction profiles. For the sake of simplicity, we assume that the Dirichlet distribution is symmetric in deriving the above formula.

Finally, the update formula for ϕk,w and θm,k are as follows:



In co-authorship networks, the vertices represent researchers and the edges in the network represent the collaboration relation between researchers. In this section we evaluate SSNLDA model on co-authorship networks collected from two distinct areas: computer science(CiteSeer) and nanotechnology(NanoSCI). Note that no name disambiguation has been done on either dataset.

The size of the largest connected subnetwork of CiteSeer is 249866 while the size of the largest connected subnetwork in NanoSCI is 203762. In this paper, we are only interested in discovering community structures in the two largest subnetworks.

In this paper, we explore three different types of social interaction pro le representations for social networks, namely 01-SIP, 012-SIP, and k-SIP.

In the 01-SIP approach, an edge is drawn between a pair of scientists if they coauthored one or more articles. Collaborating multiple times does not make a difference in this model

In order to mitigate this problem, we propose a 012-SIP model which takes a node's neighbors' neighbors into consideration. ... In this model, we distinguish a node's direct neighbors from its neighbors' neighbors by giving different weights to them.

This section describes a K-SIP model where the weight information for an edge is defined as the times of the collaboration between the two authors.

And then, 10% of the original datasets is held out as test set and we run the Gibbs sampling process on the training set for i iteration. In particular, in generating the exemplary communities, we set the number of the communities as 50, the iteration times i as 1000. In perplexity computation, i is set as 300 in order to shorten the computation time. In both case, α is set as 1/K and β is set as 0.01, where K is the number of the communities.

Table III shows 6 exemplary communities from a 50-community solution for the CiteSeer dataset with social interaction profiles being created using 012-SIP representation. ...These exemplary communities give us some flavor on the communities that can be discovered by this approach. Specifically, we observe that some communities are “institution-based”, some others are “topic-based. ... This observation reveals the fact that researchers from same institution or with similar research interests tend to collaborate together more and build closer social ties.

Perplexity is is a common criterion for measuring the performance of statistical models in information theory. It indicates the uncertainty in predicting the occurrence of a particular social interaction given the parameter settings, and hence it reflects the ability of a model to generalize unseen data.

Perplexity PP is defined as

where wm is the social interaction profiles in the test set and

where n(v)m is the number of times term t has been observed in document m.

It shows that the perplexity value is high initially and decreases when the number of communities increases. In addition, the results show that the 012-SIP approach has lower perplexity value than the other two approaches.

Compactness of a community is measured through the average shortest distance among the top-ranked Nr researchers in this community. Short average distance indicates a compact community.

The t-test results show that the 012-SIP approach is significantly better the other two approaches for both datasets

The probabilities that can be derived from SSN-LDA model can be helpful in determining the importance and roles of community members. For instance, the importance of community members conditioned on the community variable ιj can be measured through the probability p(wi|ιj ), which can be easily derived from this model based on the learned ϕ.

SSN-LDA model provides an elegant way to measure the similarity of two communities by calculating the corresponding KL (Kullback-Leibler) distance and convert it to similarity measure. KL divergence is a distance measure for two distributions and the corresponding formula for calculating the distance between two communities ιi and ιj is:



This paper describes an LDA (latent Dirichlet Allocation)-based hierarchical Bayesian algorithm, namely SSN-LDA(Simple Social Network LDA). In SSN-LDA, communities are modeled as latent variables in the graphical models and defined as distributions over social actor space. The advantage of SSN-LDA is that it only requires topological information as input.