顯示具有 co-author analysis 標籤的文章。 顯示所有文章
顯示具有 co-author analysis 標籤的文章。 顯示所有文章

2015年12月18日 星期五

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

Kucher, K., & Kerren, A. (2015). Text Visualization Techniques: Taxonomy, Visual Survey, and Community Insights. In 8th IEEE Pacific Visualization Symposium (PacificVis' 15), Hangzhou, China (pp. 117-121). IEEE Computer Society.

近年來由於可以取得大量而多樣的文本資料和採用文本處理演算法等原因,研究人員對文本視覺化(text visualization)與視覺性的文本解析(visual text analytics)的研究興趣增加。本研究針對文本視覺化技術提出一個互動的視覺調查(visual survey)。並且利用此次調查的資料,分析文本視覺化的現況,比較研究使用的各種分析與視覺化技術,以及分析有關研究者的資訊,以提供搜尋相關研究、探索次領域(subfield)以及獲得研究趨勢的洞察等目的

本研究採納前人的研究,將文本視覺化技術,以分析任務(analytic tasks)、視覺化任務(visualization tasks)、資料領域(data domain)以及資料來源(data source)、資料性質(data property)、視覺化的維度(visualization dimensionality)、視覺化的呈現(visualization representation)、視覺化的排列方式(visualization alignment)等面向,建立分類架構(taxonomy)。


分析任務是指使用者採用文本視覺化技術預期達到的主要目的,這些分類包括:
1. 文本摘要 (Text Summarization) / 主題分析 (Topic Analysis) / 實體抽取 (Entity Extraction)
2. 言談分析 (Discourse Analysis):文本或對話轉錄(conversation transcript)裡流動的語言學分析。
3. 情感分析 (Sentiment Analysis)
4. 事件分析 (Event Analysis)
5. 趨勢分析 (Trend Analysis) / 樣式分析 (Pattern Analysis)
6. 詞法/語法分析 (Lexical / Syntactical Analysis)
7. 關係/連結分析 (Relation / Connection Analysis)
8. 翻譯/文本比對分析 (Translation / Text Alignment Analysis)

視覺化任務則是由文本視覺化技術所支援的較基層呈現與互動任務,包括:

1. 自動凸顯/建議興趣區 (Region of Interest)
2. 群集 (Clustering) / 分類 (Classification / Categorization)
3. 比較 (Comparison)
4. 概觀 (Overview)
5. 監視 (Monitoring)
6. 瀏覽 (Navigation) / 探索 (Exploration)
7. 對於不確定的對策 (Uncertainty Tackling)

資料領域,包括

1. 線上社交媒體 (Online Social media)
2. 通訊 (Communication)
3. 專利 (Patents)
4. 評論 (Reviews) / 病歷 (Medical Records)
5. 文學作品 (Literature) / 詩 (Poems)
6. 科學文章 (Scientific Articles) / 論文 (Papers)
7. 社論媒體 (Editorial Media)

資料來源有單一文件 (Document) [33]、語料庫 (Corpora) [25]以及 串流文本 (Streams) [19];特殊的資料性質包括地理空間 (Geospatial) [11]、時間序列 (Timeseries) [14] 以及網路 (Networks) [6];視覺化的再現包括下列項目:折線圖 (Line Plot) / 河流圖 (River) [9, 18]、像素 (Pixel) / 面積 (Area) / 矩陣 (Matrix) [13, 7, 4]、節點-連結 (Node-Link) [32]、雲 (Clouds) / 銀河 (Galaxies) [1, 3]、地圖 (Maps) [34]、文本 (Text) [26]與形符 (Glyph) / 圖標 (Icon) [28, 10];排列則包括了輻射狀 (Radial) [35]、線性 (Linear) / 平行線 (Parallel) [8] 以及測標依賴 (Metric-dependent) [22]。

本研究指出有超過一半(56%)的文本視覺化利用主題模型(topic modeling)技術,資料來源方面大多數支援語料庫(70%),並且許多支援時間相關的資料(43%),而視覺再現方面主題以二維(2-D)為主,僅有極少數的研究以三維(3-D)的方式呈現,約占所有研究的4%。

文本視覺化的前五位主要作者為Daniel A. Keim (17 筆)、Shixia Liu (12 筆)、Christian Rohrdantz (9 筆)、Daniela Oelke (7 筆)和 Huamin Qu (7 筆)。將作者依據他們的合著關係建立研究者合作網路圖後,觀察網路圖的相連成分,可以發現大部分是獨立的小群體,最大的成分上共有106位作者,並且在這個成分上的兩個主要集群為University of Konstanz和Microsoft Research Asia等兩個研究團隊,Daniel A. Keim 和 Shixia Liu分別為集群的中心,並且他們二位也是網路圖上中介中心性最高的節點。雖然在本研究蒐集的資料上,這兩位作者之間並沒有直接的合作關係,但他們都曾與中介中心性第三高的兩位作者Dongning Luo 和 Jing Yang合作。


In this paper, we present an interactive visual survey of text visualization techniques that can be used for the purposes of search for related work, introduction to the subfield and gaining insight into research trends.

The interest for text visualization and visual text analytics has been increasing for the last ten years. The reasons for this development are manifold, but for sure the availability of large amounts of heterogeneous text data (caused by the popularity of online social media) and the adoption of text processing algorithms (e.g., for topic modeling) by the InfoVis and Visual Analytics communities are two possible explanations.




Analytic Tasks
these items are critical to the main analysis goals that users expect to achieve when employing a text visualization technique.

1. Text Summarization / Topic Analysis / Entity Extraction

2. Discourse Analysis
the linguistic analysis of the flow of text or conversation transcript.

3. Sentiment Analysis
for techniques related to the analysis of sentiment, opinion, and affection.

4. Event Analysis
deal with the extraction of events from the text data or involve visualization of text in some different manner

5. Trend Analysis / Pattern Analysis
both automated trend analysis and manual investigation directed at discovering patterns in the textual data.

6. Lexical / Syntactical Analysis

7. Relation / Connection Analysis

8. Translation / Text Alignment Analysis

Visualization Tasks
lower-level representation and interaction tasks that are supported by the text visualization techniques.

1. Region of Interest
the automatic highlighting/suggestion of data items/regions that could be of interest to the user for more detailed investigation

2. Clustering / Classification / Categorization

3. Comparison

4. Overview
both techniques that provide “the big picture” by displaying a significant portion of the data set as well as techniques which use special aggregated representations to provide overview while reducing the visual complexity

5. Monitoring

6. Navigation / Exploration

7. Uncertainty Tackling


Domain

1. Online Social media

2. Communication

3. Patents

4. Reviews / (Medical) Records

5. Literature / Poems

6. Scientific Articles / Papers

7. Editorial Media

Data sources include the following self-evident items: Document [33], Corpora [25], and Streams [19].

The special data properties include Geospatial [11], Timeseries [14], and Networks [6].

Representation includes the following items: Line Plot / River [9, 18], Pixel / Area / Matrix [13, 7, 4], Node-Link [32], Clouds / Galaxies [1, 3], Maps [34], Text [26], and Glyph / Icon [28, 10].

Alignment, i.e., layout, includes Radial [35], Linear / Parallel [8], and Metric-dependent [22].

As displayed in the table, our proposed taxonomy includes most of the categories except for two: we believe that the underlying data representation (e.g., bag-of-words vs. language model [30] or whole text vs. partial text [24]) is more relevant to the underlying computational methods than to observable visualization techniques.

And the same naturally holds for data processing methods (e.g., the specification of involved MDS methods [2]) that are partially covered by other categories in our taxonomy, for instance, the analytic task of topic analysis implies the usage of corresponding computational methods.

Using the data collected for the survey, we have been able to analyze the general state of the text visualization field, to compare the usage of various analysis and visualization techniques (with regard to our taxonomy), and to analyze the information about researchers in this field.

According to our current set of entries, the trend for rapid increase of text visualization techniques started around 2007.

With regard to category statistics (cf. Fig. 4), there is an obvious interest for tasks related to topic modeling (56% of all entries).

The majority of the techniques support corpora as data sources (70% of all entries), and a lot of them support time-dependent data (43% of all entries).

Another result—which is probably expected—is that only less than 4% of all entries use 3-dimensional visual representations.

We have also taken a look at the authorship statistics for the current data set. The top five authors with regard to number of techniques are Daniel A. Keim (17 entries), Shixia Liu (12 entries), Christian Rohrdantz (9 entries), Daniela Oelke (7 entries), and Huamin Qu (7 entries).

As seen in Fig. 5, the majority of author nodes are included into isolated connected components of small sizes (less than 10 nodes) while there is a big connected component with 106 nodes present in the graph.

The two major clusters in that component represent the research groups from the University of Konstanz and Microsoft Research Asia with Daniel A. Keim and Shixia Liu as cluster center nodes.

Shixia Liu and Daniel A. Keim happen to have the 1st and the 2nd largest betweenness values in the graph, respectively. While these two researchers have no direct collaboration with regard to our data set, they both have collaborated with Dongning Luo and Jing Yang who both share the 3rd largest betweenness value.

2014年2月15日 星期六

Wouters, P., & Leydesdorff, L. (1994). Has Price's dream come true: Is scientometrics a hard science?. Scientometrics, 31(2), 193-222.

Wouters, P., & Leydesdorff, L. (1994). Has Price's dream come true: Is scientometrics a hard science?. Scientometrics, 31(2), 193-222.

本研究利用Scientometrics期刊論文以及其參考文獻為研究資料,根據多種資訊判斷科學計量學領域是否已經是硬科學,並分析這個領域的其他特性。本研究所使用的科學計量學資訊有引用文獻的相對年齡(relative age of the cited literature)、論文作者間的關係、論文題名的詞語模式(patterns of words  in the titles of these articles)等。

根據Price的知識增長理論(theory of knowledge growth),科學家會引用本身領域的文獻,因此,如果有研究前沿(research fronts)存在於這個領域,便會產生立即效應。Price指標(Price index)可以測量立即效應(immediacy effect),Price指標較大表示引用文獻的相對年齡較低,例如Price(1970)測得生物化學和物理的Price指標值約在60%到70%,社會科學大約在42%附近。Crane(1972)則認為科學會形成作者間彼此緊密相連的社群,因此本研究分析作者間的合著關係和引用的關係,並且利用網絡分析技術探討科學社群的凝聚程度,並且測量作者在網絡的位置連結性以及結構的相似性,根據這些資訊進行叢集,集結彼此間連結性強的作者形成一個叢集,或是形成位置相似的作者叢集。另外,Rip and Courtial (1984)和Leydesdorff (1989a)
指出題名上的詞語可視為是出版品的認知訊息(cognitive message)的指標,詞語在題名上的共現可視為是詞語間關係存在的紀錄,因此本研究也利用網絡分析技術探討詞語的共現網絡。

研究結果發現分析的779筆Scientometrics期刊論文資料,除了前三年快速的增加外,平均每年增加3.5筆,並且這些論文資料共包含12341筆參考文獻,平均每篇論文有15.8筆參考文獻。各項指標都相當穩定。Price指標的平均值為43.0%,若以每年的Price指標的平均值在34.0%到51.4%之間。

以作者資料來看, 779筆論文資料共計由669位不同的作者完成,有接近3/4的作者(488位)僅出現在一筆論文資料上,每位作者平均出現在1.8筆論文資料上,其作者生產力符合Lotka分布,並且大部分(61%)的論文是單一作者,平均每篇論文有1.6位作者。合著作者的論文資料中,大多數的作者都僅和一到兩位同事合作,合著網絡相當破散,但幾個較大的網絡與作者在同一機構任職、參與同一研究計畫或者具有共同的研究興趣有關。

Scientometrics期刊論文的作者引用網絡則呈現高度凝聚的狀態。779筆論文資料中有441筆被其他Scientometrics期刊論文引用,每筆Scientometrics期刊上的論文引用的論文平均有19.4%同樣是Scientometrics期刊的論文。發表超過1篇以上論文的作者共有181位,其中的130位作者有引用其他129位作者的資料,利用作者之間的彼此互相引用關係,發現形成的集團(clique)大多與作者任職機構有關,也有一個集團是成員間曾彼此辯論(debate)而產生。最後,題名上的詞語共現網絡也同樣有高度凝聚的情形。從上面的資訊可以判斷科學計量學領域已經由多種的學科背景在認知與社會性上整合而成,但並沒有發現研究前沿的現象。

In more than one respect, Scientometrics displays the characteristics of a social science journal. Its Price Index amounts to 43.0 percent, and is remarkably stable over time.

The majority of the published items in Scientometrics has been written by a single author. Moreover, the network of co-authorships is highly fragmented: most authors cooperate with no more than one or two colleagues.

Both the citation networks of the authors and the network of title words indicate that the field is nonetheless highly cohesive.

The characteristics of the publications in this journal, and the patterns of the bibliometric relations among them, may therefore indicate the type and extent of the cognitive and social integration of the various disciplinary backgrounds into scientometrics as a field.

The question, in other words, is how "hard" scientometrics is, and how strongly its knowledge is codified. These properties can be measured in terms of:
a) the relative age of the cited literature, the so-called "Price Index",
b) the relations among the authors of articles published in Scientometrics; and
c) the pattern of words in the titles of these articles.

According to Price's theory of knowledge growth (Price 1965), science distinguishes itself from other fields of study by the way scientists refer to their literature (Price 1970). The existence of "research fronts" in science supposedly leads to an "immediacy effect", which can be measured in terms of the so-called "Price Index".

The Price Index is defined as "the proportion of the references that are to the last five years of literature" (Price 1970). Price estimated that this index would vary between 22 and 39 percent if no immediacy effect were present. [1] A field that was all research front and with no general archive might have a Price Index of 75 to 80 percent.

From his analysis of 162 journals, Price (1970) concluded: "Perhaps the most important finding I have to offer is that the hierarchy of Price's Index seems to correspond very well with what we intuit as hard science, soft science, and nonscience as we descend the scale." Biochemistry and physics are at the top, with indexes of 60 to 70 percent, the social sciences cluster around 42 percent, and the humanities fall in the range of 10 to 30 percent.

Science is, on the whole, practised in tightly knit communities in which the authors address one another (Crane 1972).

Co-authorship relations can be considered as indicators of co-operation. [3]

The meaning of citation relations is less clear, given the ongoing citation debate (MacRoberts and MacRoberts 1989; Cozzens 1989; Luukkonen 1990; Leydesdorff and Amsterdamska 1990; Woolgar 1991). But whatever the precise meanings of citations may be, citations can be considered as sociometric data, and the resulting network can accordingly be analyzed (cf. Shrum and Mullins 1988).

We analyzed the extent to which the authors are connected to one another, i.e. the cohesiveness of the network, as well as the pattern displayed by each author in relation to all other authors, i.e. the position of authors in the network.

We also analyzed the similarities among authors in both these dimensions of the matrices, i.e. we clustered strongly connected authors as well as authors in similar positions. Direct as well as indirect linkages between the authors are involved in this analysis.

Strong cliques are sets of authors connected by relations in such a way that all members of the clique are connected to one another, and anyone for whom this holds is included in the clique. The inclusion criterion is less strong for weak cliques, in which all pairs within the clique must have relationships with all other pairs, and anyone with a relation to or from a member of the clique is included.

Strong structural equivalence clusters are sets of authors with completely identical positions in the network (the distance dP between them is zero). Weak structural clusters are sets of authors with a significant similarity in their patterns of relations (the distance dP is small).

As noted, we wished to know whether a structurally codified semantics of scientometrics exists or whether, on the contrary, the articles in Scientometrics use the different terminologies of the various disciplines surrounding scientometrics.

Given the functions of titles of articles, the words in these titles can be considered as indicators of the cognitive message of the publication (Rip and Courtial 1984; Leydesdorff 1989a). The co-occurrence of words in titles can be considered as an indication of the existence or non-existence of relations between these words (Callon et al. 1983).

Since 1978, 779 items have been published in Scientometrics. They contain 12,341 references to the scientific literature. [10]

The number of publications per year in the journal increases in a linear way (Fig. 1). [11] After a steep growth during the first three years, the number increases by 3.5 publications per year.

The number of references per year shows a comparable pattern, although somewhat more irregular. Every publication contains on average 15.8 references (cf. Yitzhala, 1991). Since 1986, this number has become stable at an average of 15 references per publication (Fig. 2).

The publications in Scientometrics were written by 669 different authors. On average, every author published 1.8 times and every paper was written by 1.6 authors. Nearly three-fourth of the authors (488 or 73 percent) published only once in Scientometrics. The distribution of productivity among the authors is a Lotka distribution (Fig. 4).

The average Price Index of Scientometrics is 43.0 percent. ... The Price Index varies between 34.0 and 51.4 percent (Fig. 5). The regression line is not significant. [12] Apparently, the index displays neither rise nor fall since 1978.

Recently, Schubert and Maczelka (1993) concluded from an analysis of Scientomettics in 1980-81 and 1990-91 that the journal has moved slightly from the "soft" (social) towards the "harder" (natural) sciences. They drew this conclusion from the rise of the Price Index from 35 percent to 42 percent between these measurement points. This observation is, however, based on only two measurements. Because of the statistical fluctuations in the value of the Price Index over time, any conclusion can be drawn regarding the development of the Price Index if one restricts oneself to only two measurement points.

In accordance with Price's theory, the number of references to literature of a specific age rises until the cited literature is two years older than the citing literature, and then falls off (Fig. 7). Note that this decline is gradual. Apparently, only a small "immediacy effect" is visible in scientometrics.

A general phenomenon in science is the growth of the number of co-authored scientific articles, relative to the total scientific production (Luukkonen et al. 1992; Abt 1992).

In Scientometrics, however, 61 percent of the articles have been written by a single author. This share is stable over time.

The network of co-authorships is highly fragmented. ... With the exception of three subgroups, most co-authors cooperate with no more than one or two colleagues.

Comparison of the composition of the weak structural equivalence clusters with the relational cliques reveals that two clusters are identical: a group of authors from Leiden (Van Raan et al.) and a group of authors with various institutional affiliations, probably best characterized as the "co-word analysis group". So, these two groups have distinct identities, with respect both to their relations and to their positions in the network.

Some clusters seem constituted by the institutional affiliations of the authors. This holds for the Leiden group and for the authors around ISI (cluster 3). In other cases, nationality appears to be the binding force. This holds for the group in Hungary (cluster 5), the Belgian informetricians (cluster 10) and the Spanish scientometricians (cluster 6). However, cluster 1 can best be characterized by its research program (co-word analysis). Cluster 2 seems to consist of authors from Sussex together with CHI Research Inc. Thus, co-author relations are not only institutionally defined; shared interests and common intellectual goals play a role as well.

To sum up, scientometrics is a fragmentary field of co-authorships. The authors are highly selective in their co-authorship relations with one another. Co-authorships are defined neither exclusively by social nor only by intellectual factors. Both dimensions shape the pattern of co-authorships.

With respect to the number of solitary authors and the large number of isolated small clusters, scientometrics exhibits the pattern of a social science.

Of the 779 articles published in Scientometrics, 411 were subsequently cited one or more times in Scientometrics. The share of references to Scientometrics (as a percentage of all references) has stabilized around an average of 19.4 percent since 1987.

Of the 181 authors in the core set, 130 authors cite one another. So, 51 (or 28.2 percent) of the authors publishing more than one article in Scientometrics from 1978 till 1993 are neither citing nor cited within this group of authors.

The core set of authors in Scientometrics is found to be highly cohesive in terms of their mutual citation relations. All these authors are members of one single weak clique. Moreover, a majority of these authors (88) also belongs to one strong clique (Table 5).

The picture is different if we exclude all indirect relations from the analysis. This "fine structure" of the citation matrix is shown in Table 6, where 13 strong cliques and 6 weak cliques are revealed. Most strong cliques seem to coincide with shared institutional affiliations. The exception is clique 9, which indicates the existence of a debate among the members of this clique.

The most striking feature of the network of title words of articles published in Scientometrics is its cohesiveness. All words cluster together in a single strong component clique (Table 8). If only direct relations are included, all words cluster together in a single weak component clique. This means that all words are either used together in a title or share a common co-word.

Thus, the language of scientometrics is both strongly unified and weakly codified. This strong cohesiveness is a stable characteristic of the titles in Scientornetrics, from the very start of the journal. Perhaps a distinct discourse already existed before the journal was founded. In any case, it constitutes a textual identity of scientometrics as a field, one probably different from the various mother disciplines. Thus a process of de-differentiation seems to have occurred not only in the patterns of citing (and being cited) but also at the cognitive level.

The interpretation of the Price Index is complicated because of these variations within disciplines. If we, nevertheless, take the Price Index preliminary as an indicator of "hardness", scientometrics belongs to the group of relatively hard social sciences. At the same time, it stays unequivocally within the social science range. Taken literally, Price's dream has therefore not come true, since he postulated the emergence of a completely new type of social science with a natural science character. But if we reformulate his goal a posteriori in a more modest way, as the building of a relatively hard social science, it did come true.

The value of the Price Index appears stable over the years. Since a number of other indicators also exhibit stability, this seems to suggest the existence of some scientometric identity. For example, the journal expands at a regular rate, while the percentage of co-authored papers increases only very slowly. The origin of this stability can best be explained by the finding that the community of researchers who have published more than once in Scientometrics acts as a tightly knit network.

In addition to the co-authorships within various institutes, and partly overlapping with this structures, there are national co-authorship relations, like those among the Belgian informetricians, and programmatic co-authorship relations, like those among the users of the French co-word instrument. In general, co-authorship relations are firmly embedded in existing social structures, both at the national and at the community level.

These various strong graphs of co-authors, however, are structurally embedded in the communication structure as indicated by textual indicators. Both in terms of citation relations and in terms of title-words the network is very cohesive, while the structural dimensions of codification are less clear.

In summary, the community of authors publishing in Scientometrics is well integrated, while there are no indications of an exclusive paradigm or a research front.

2014年1月24日 星期五

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

本研究探討在地理上的生產力遷移(shifts in productivity)是否發生在書目計量學(bibliometrics)、資訊計量學(informetrics)和科學計量學(scientometrics)等計量學(metrics)領域,也就是歐洲的貢獻明顯地成長,並且北美的貢獻相對來說有減少的情形。

有關計量學的研究,Hood and Wilson (2001)和Stock and Weber(2006)等研究都分析了這個領域的文獻成長情形。Hood and Wilson (2001)回顧了計量學領域的發展,並且比較bibliometrics、scientometrics和informetrics的相關文獻,發現bibliometrics還是在相關領域上使用最廣泛的詞語。Stock and Weber(2006)從觀察中確認這個領域從1980年後便持續地成長。Wolfram (2008)則發現在計量學領域中,北美的文獻有明顯地減少而歐洲則是急遽地增加的情形。

本研究利用bibliometrics、scientometrics、informetrics、cybermetrics、webometrics、citation analysis、link analysis和citation indexes做為檢索的問句,同時再加上Scientometrics和Journal of Informetrics兩種期刊的論文,從Web of Science資料庫中進行檢索。結果共檢索出4404筆論文資料。

在這些論文資料裡,共有75個國家。以地區來區分,歐洲在每個時段上具有最大的貢獻,不論是數量或所占比率都有成長,亞洲所佔的相對比例在22年間有很大的成長,北美雖然在數量上有成長,可是相對的比例呈現緩慢的下降。每個地區的作者會偏好在本身地區的期刊上發表,舉例而言,歐洲作者發表論文的前五個期刊中有四個歐洲期刊,南美也有類似的情形,但是亞洲的情形例外,前五個期刊中有四個是歐洲期刊,另一個則是北美的期刊。

自1990年代中期後,國家間的合作情形增加許多,之前國際合作的論文每年為1到19篇,2008年已大幅增加為96篇。美國是國際合作佔最多的國家,但以地區來說,歐洲平均每個國家的國際合作數為5.78篇論文,多於世界其他部分的4.47篇論文。

此外,歐洲則有許多具有國際合作經驗的機構,共有16所研究機構有國際合作經驗,北美則有8所,亞洲有1所。機構間的合作來說,在1987年每篇論文平均只有1.1個機構,但在2007年則增加為1.96。

本研究且利用MDS、VOSviewer和Pajek將這些論文上的國家與機構之間的合作關係,呈現為圖形。

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).


This investigation was prompted by interest in whether shifts in productivity based on geography are observed in the bibliometrics, informetrics and scientometrics areas.

One of the authors conducted a pilot study to determine whether there have been clear declines in North American contributions to the metrics literature base (Wolfram, 2008). The author found that there was indeed a notable relative decline in North American contributions and a sharp increase in European contributions.

Hood and Wilson (2001) examined the growth of literature of the metrics area. They provided an historical treatment of the development of these areas that included earlier studies of the field. In their research, literature associated with bibliometrics, informetrics and scientometrics was compared for the period 1968–2000. The authors noted that bibliometrics was still the most widely used term for metrics research.

More recently, Stock and Weber(2006) conducted a Web of Science search for records specifically including metrics terms and allied areas. They observed contributions had grown substantially since 1980.

Search parameters included the Boolean ORed result of bibliometrics, scientometrics, informetrics, cybermetrics and webometrics, in truncated form (e.g., webometri*), along with the phrases “citation analysis”, “link analysis” and “citation indexes”. ... These search results were ORed with the two primary journals that publish metrics research that are indexed by WoS, namely Scientometrics and the Journal of Informetrics.

A pair-wise comparison of all collaborations at the national and institutional levels was then conducted from which a cooccurrence matrix could be compiled.

Multidimensional scaling (MDS) analysis was used to visualize the relationships among countries. Because the data represent a type of similarity measure represented as a symmetric matrix, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006).

The recently developed visualization tool VOSviewer (van Eck &Waltman, 2010) was also used to provide an alternate visualization of the relationship outcomes. Like MDS, VOSviewer (http://www.vosviewer.com/) relies on a distance-based approach to mapping informetric relationships. Instead of using more traditional similarity measures to produce a normalized outcome for co-occurrences as used in MDS, relationships are based on association strengths, so the algorithm is somewhat different than PROXSCAL and, therefore, can produce different outcomes. Details of the comparison of different measures can be found in van Eck and Waltman (2009).

The network visualization software Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/) was used as well. Unlike the distance-based mapping of PROXSCAL and VOSviewer, Pajek produces directed or undirected network maps, with the strength of the relationships represented by the thickness of connecting lines between vertices on the map. Distances are used more for clarification, but proximities do not necessarily indicate a stronger relationship.

The search parameters retrieved 4404 publications.

Europe shows the highest levels of contribution, both in absolute and relative terms over the time period of the study. Growth patterns in absolute terms are nonlinear based on trend line analysis in MS Excel; however, the R-squared goodness-of-fit values for even the best fitting models (higher order polynomials) were never more than 0.95, indicating a less than desirable fit.

Relative contributions based on geographic divisions have been largely stable. An exception is Asia, which had an increasing relative contribution over the 22-year time frame of the study. Although North American contributions have continued to increase in absolute numbers, the relative contribution shows a slow average decline over time.

The top five journals listed for each continent demonstrated a regional preference for publication outlets from that region. So, for example, four of the top five journals for European publications were published in Europe, and four of the top five journal outlets for South America were South American. The exception to this was Asia. Four of the top five journals for Asian publications were European and one was North American. This outcome may be a reflection of the data extraction method, the indexing practices of WoS, or a preference during the study time frame for Asian scholars to publish in Western journals.

Seventy-five countries were represented in the record set.

The number of metrics papers published annually that represent collaborations between two or more countries has increased greatly since the mid-1990s. Prior to this time, the number of internationally collaborative papers ranged from 1 to 19 papers annually. Over the last decade this number has increased to a high of 96 papers in 2008.

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).

Sixteen of the institutions on the list are European, eight are North American, and one is Asian. The United States has the largest number of institutions represented (five), followed by Belgium (four – note: one institution merged with another institution to form a new entity).

There has been steady growth in inter-institutional collaboration over the 22 years. The mean number of collaborative institutional partners within the dataset has steadily increased from a low mean of 1.1 institutions per publication in 1987 to a high of 1.96 institutions per publication in 2007.

Europe, and in particular Western Europe, clearly dominates in the production of metrics literature. The United States continues to be the largest singular contributor, but this appears to be changing. North American contributions as a whole continue to increase, but represent a smaller percentage of worldwide production. European contributions have grown tremendously, especially during the last 5 years of the study period. This same period is marked by impressive growth from Asia.

It should be noted that WoS increased its coverage in 2008 by including more regional journals. These inclusions possibly could contribute to the increase in Asian contributions, but the observed growth for Asia was already evident prior to any such additions.

International and inter-institutional collaborations do not necessarily reveal strong geographic affinities, although the multiple institutional affiliations by a number of scholars associated with Flemish institutions do contribute to the strengthening of regional ties. Undoubtedly, the growth of the Internet and increasing availability of other telecommunication technologies have made these collaborations less distance dependent.

2014年1月18日 星期六

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

Hou, H., Kretschmer, H., & Liu, Z. (2008). The structure of scientific collaboration networks in Scientometrics. Scientometrics, 75(2), 189-202.

本研究利用社會網絡分析、共現分析(co-occurrence analysis)、叢集分析和詞語的頻率分析等多種分析技術,從Scientometrics期刊1978到2004年發表的1927筆論文資料,探討科學家合作網絡的結構特性、整個網絡上的合作領域以及個別的合作網絡、合作網絡上的合作中心(collaborative  center)。

過去的研究裡,Schubert (2002) 和 Dutt, Garg, & Bali (2003)都是針對國家間合作的巨觀層次。Kretschmer (2004) 認為巨觀和中觀(meso)層次的分析無法足夠地反映個人之間的合作趨勢,因此呼籲應在微觀層次的分析投注更多努力。

1927筆論文資料裡,單一作者的論文共有1052筆,所以仍稍占多數。作者數大於3的論文僅占非單一作者論文的13.71% (120/875),顯然研究Scientometrics的團隊規模都不大。發表3篇論文以及以上的高生產作者共計234人,其中有69.66%的作者曾發表與其他作者合作的論文。將這些作者間的合作關係表現成網絡,並利用Bibexcel對這個網絡上的節點進行叢集分析,共發現22個叢集。前兩個較大的叢集分別有15與14個科學家。網絡上最大的相連成分上共有15個叢集,共有合作經驗的高生產作者中的96位,占58.90%。合作網絡共有401條連結線,網絡密度為0.03,顯示Scientometrics領域的合作很鬆散。

對每一個節點計算它們的三種中心性,結果發現中心性和對應作者的生產力之間有很顯著的正相關,表示高生產力的作者同時也活躍在Scientometrics領域的合作網絡上。其中Glänzel的程度中心性最高,總共和其他18位作者有合作關係。

以詞語的頻率分析每個叢集的主題,最大的兩個叢集有類似的主題,但使用的研究方法略有不同。此外,研究主題為科學合作的四個叢集間幾乎沒有連結,同樣的情形也發生在研究科學與技術之間關係的四個叢集。

The structure of scientific collaboration networks in scientometrics is investigated at the level of individuals by using bibliographic data of all papers published in the international journal Scientometrics retrieved from the Science Citation Index (SCI) of the years 1978–2004.

Combined analysis of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of words is explored to reveal: (1) The microstructure of the collaboration network on scientists’ aspects of scientometrics; (2) The major collaborative fields of the whole network and of different collaborative sub-networks; (3) The collaborative center of the collaboration network in scientometrics.

Schubert [8] and Dutt etc. [9] presented international collaboration characteristics in the scientometrics community itself, focusing on country aspects at macro level.

Kretschmer [6] appealed to devote more efforts to investigations at micro level in the future because the knowledge at meso and macro level does not yet adequately reflect the trends in cooperation between individuals.

The study is based on bibliographic data retrieved from the Web of Science. The data contains all types of documents published in Scientometrics during 1978 to 2004.

In this study we have adapted an integrated procedure of social network analysis (SNA), co-occurrence analysis, cluster analysis and frequency analysis of title words.

Bibexcel is designed as a tool for manipulating bibliographic data, which is a free online-software published by Persson. In the present study, Bibexcel is used to do cooccurrence analysis and cluster analysis.

Following the methods of Otte & Rousseau [11], White [13] and Kretschmer & Aguillo [12], SNA was applied to display the microstructure of collaboration networks in scientometrics with Pajek.

Moreover, we used frequency analysis of title words to display the main collaborative field of different sub-networks. The software for frequency analysis is demo version of Wordsmith Tools published by Oxford University Press and available online.

There were 1927 documents published in Scientometrics during 1978 to 2004 (see Table 1).



From Table 1, we found that the pattern of co-authorship was still dominated by single-authored papers as the conclusion drawn by Dutt etc. [9].

While the number of multi-authored papers (the number of co-authors is more than 3) accounts for 13.71% only, which indicates that team size in scientometrics is not large.

In order to show the main structure of the network, each author must published 3 papers or more to be included in this integrated analysis. This threshold resulted in a total of 234 prolific authors publishing 3 or more papers during 1978 to 2004, among them there are 163 authors published co-authorship papers, accounting for 69.66% of the prolific authors.



Based on cluster analysis embedded in Bibexcel, we gained 22 clusters circled by solid lines (see Figure 1). We identified these clusters as sub-networks in the field of scientometrics.

The largest subnetwork is number 1 that has 15 collaborators, and the second largest one is number 2, which has 14 collaborators, and so on.

We noticed that there was totally 15 subnetworks connected with each other composing the largest central component, which had 96 numbers accounting for 58.90% of the prolific authors published co-authorship papers.

Density is an indicator for the general level of connectedness of the graph. ... In the present study, there are totally 401 links in the network, so the density of the network is 0.03, which indicates that the collaborative network in the field of scientometrics is very loose.

So an author who has high degree centrality must has collaborated with many other authors, which means the author is a central collaborator of the whole network. In the present study, Glänzel who has 18 co-workers is the central author of the whole network.

We found a positive and significant correlation between output of authors and the centrality measures (r=0.648, 0.437, 0.338 respectively at the 0.01 level, see Table 4) after investigating the correlations between output and the three centralities of the 125 authors in the 22 sub-networks, which indicated that most of the prolific authors are also active in collaboration network in the field of scientometrics.

We have also presented the main collaborative field of different sub-networks in scientometrics and found that the two biggest sub-networks have the similar collaborative topic with slightly methodological difference. In addition, we found an interesting phenomenon that four sub-networks dealing with scientific collaboration didn't collaborate with each other except sub-network 3 and 12. Moreover, four subnetworks studying technology and science never collaborated with each other at all.

2014年1月17日 星期五

Chen, Y. W., Fang, S., & Börner, K. (2011). Mapping the development of scientometrics: 2002–2008. Journal of Library Science in China, 3, 131-146.

Chen, Y. W., Fang, S., & Börner, K. (2011). Mapping the development of scientometrics: 2002–2008. Journal of Library Science in China, 3, 131-146.

本研究利用社會網絡分析與科學地圖映射(science mapping)分析Scientometrics期刊2002到2008年發表的816筆論文。

針對Scientometrics期刊進行書目計量分析的相關研究,包括:Schoepflin and Glanzel (2001)將Scientometrics在1980、1989和1997年發表的論文分別進行歸類,發現科學政策(science policy)和科學社會學(the sociology of science)的比率在下降。Peritz and Bar-Ilan (2002)發現Research Policy和Social Studies of Science分別是1990和2000年Scientometrics論文引用的期刊次數最多的第三名和第四名。Chen, McCain, White, and Lin (2002)分析出1981到2001年間Scientometrics期刊的引用及共被引模式。Hou, Kretschmer, and Liu (2008)對2002到2004年間Scientometrics期刊上的作者合作網絡的結構特性進行分析。Dutt, Garg, and Bali (2003) 則分析Scientometrics期刊1978到2001年間論文資料上的國家、機構在主題上的分布。

本研究在816筆論文資料上共計發現57個國家,具有較大生產力的國家主要是歐洲國家。前十個較大生產力的國家裡,美國、比利時、西班牙、中國和德國都有相當快速的年增率,但印度的年增率是負的。生產力較大的國家的被引用次數也比較高。

為了國家間的研究合作情形,本研究提出相對合作強度(relative collaborative intensity, RCI),這個測量方式整合了合作的國家數和合作的次數兩種指標,其公式如(3)所示:

假設(RCI)i是第i個國家的相對合作強度,其中CCiCTi分別是這個國家合作的國家數和與其他國家合作的次數。在本研究裡,比利時是相對合作強度最高的國家,英國、荷蘭與美國則分居2到4名。

接下來將國家間的合作關係表現成網絡圖,圖形上最大的相連成分(connected component)共有37個國家。在這個相連成分上,比利時、英國和匈牙利之間都有很強的連結。

以機構來看,比利時的Katholieke Univ Leuven、匈牙利的Hungarian Academy Science和荷蘭的 Leiden Univ發表的論文數和被引用次數最多。

進一步分析前十個主要機構的被引用次數最多的前十筆論文資料,發現引用它們的論文主要來自圖書資訊學、電腦科學、資訊系統和跨領域應用(interdisciplinary applications)等領域。但台北醫學大學的一篇論文則被許多生物醫學領域的論文引用。

就論文的合作作者數來分析,本研究發現單一作者的論文有271篇,多位作者的論文有545篇,每篇論文平均有2.29位作者。Dutt, Garg, and Bali (2003) 研究1978-2001年間的論文,單一作者的論文占半數一上,平均合作作者數則為1.73。兩相比較之下,由多位作者的論文數和平均作者數增加的結果,能夠顯示Scientometrics期刊上的合作情形增多。

從引用的文獻分析Scientometrics的主題包括科學與技術的關係(the relationship between science and technology)、個人科學研究產出的量化指標(indexes to quantify an individual's scientific research output)、作者的合作現象(author collaborations)、共被引網絡(co-citation networks)、科學引響力以及國家富強(the scientific impact and wealth of nations)。

The purpose of this article is to use the methods of Social Network Analysis and Science Mapping to make an analysis on the 816 papers published in the international journal Scientometrics from 2002 to 2008.

The major tools used in this paper were TDA, NWB and Excel.

Börner (2006) discussed the mapping research on structure and evolution of science.

Börner, Penumarthy, Meiss, and Ke (2006) mapped the diffusion of information among 500 major U.S. research institutions based on the 20-year publication data set published in the Proceedings of the National Academy of Sciences (PNAS) in the years 1982-2001.

Boyack, Börner, and Klavans (2009) mapped the structure and evolution of chemistry research over a 30 year time frame based on Science (SCIE) and Social Science (SSCI).

Leydesdorff and Rafols (2009) made a global map of science based on the ISI subject categories.

For instance, Schoepflin and Glanzel (2001) found a decrease in the percentages of both the articles related to science policy and to the sociology of science by classifying the articles published in Scientometrics in the years 1980, 1989 and 1997.

Peritz and Bar-Ilan (2002) analyzed the papers published in Scientometrics in 1990 and 2000 and found that Research Policy and Social Studies of Science are the third and fourth most frequently referenced journals in articles published in Scientometrics.

Chen, McCain, White, and Lin (2002) drew upon citation and co-citation patterns derived from articles published in the journal Scientometrics (1981-2001).

Hou, Kretschmer, and Liu (2008) analyzed the structure of scientific collaboration networks in scientometrics at micro level (individuals) by using bibliographic data of all papers published in Scientometrics of the years 2002-2004.

Dutt, Garg, and Bali (2003) made an analysis of papers published by Scientometrics during 1978 to 2001 by scientometrics assessment on countries and themes distribution, comparison of institutions and co-authors.

The analysis of 816 papers published in Scientometrics during 2002-2008 showed that they were contributed by 57 countries (or regions). ... Most of the 57 countries were from Europe. Other major countries (or regions) had a larger number of papers were USA and Canada in North America, China, India, Taiwan, South Korea and Japan in Asia, Brasil in Latin America, and Australia.

Fig. 2 had clearly illustrated the average annual growth rates of TOP10 countries, from which we can conclude that USA, Belgium, Spain, China and Germany had higher growth rates and India had a negative growth rate.

From Fig. 3 we could see that all the TOP 10 countries had a higher number of times cited. It indicated that the papers contributed by those countries were of higher quality and had more impact.



In order to visualize the relative intensity of collaboration, this article introduced the concept of Relative Collaboration Intensity (RCI) indicator. The average number of collaboration countries (CC), average collaboration times (CT) and Relative Collaboration Intensity (RCI) of the 10 countries were given in formula (1), (2) and (3):



We found that Belgium had the highest relative collaboration intensity, and England, Netherlands and USA ranked 2, 3 and 4.

In order to make a clear vision about the collaborations among all the countries/regions (57), the country collaboration network had been made with the method of SNA by NWB. ... The largest connected component in the network had 37 nodes, and there is another small component with 2 nodes.

Fig. 4 showed the largest component with 37 countries, which depicted that Belgium, England and Hungary had formed an strong connection. The largest connection lied between Belgium and Hungary, and the collaboration times were 27. Fig. 4 also showed that although the USA had the largest number of papers, the collaboration activity was weaker than Belgium, Hungary, England and Finland. USA had paid much more attentions to collaborate with Canada, England and Australia. Netherlands had collaborated with many countries, however, the collaboration times were fewer compared to Belgium, England and Finland.

The data showed that Katholieke Univ Leuven (Belgium), Hungarian Academy Science (Hungary) and Leiden Univ (Netherland) ranked from first to third both in number of papers and times cited and all of them had a biggish advantage to others.

We select the Most-Cited paper (that had the highest value of times cited) of each TOP 10 TC/P institutions and get 10 Most-Cited papers finally. By analyzing their citing papers, we found that the citing papers which had cited the Most-Cited paper of each institution distributed mainly in the fields of information science & library science, computer science, information systems and interdisciplinary applications ....

So a conclusion could be made that although an institution did not have many papers or hold the advantage of research activities, it could carry out one or some significant works that had a great impact on the future development of information science & library science. And some research work on scientometrics had also affected the development of some other scientific fields, such as the work of Taipei Med Univ.

Another study carried out by Dutt et al. (2003) in scientometrics showed that the average number of authors per paper was 1.73 during the period of 1978-2001. We studied the average number of authors per paper published in Scientometrics 2002-2008 and found that the value was 2.29, which indicated that collaboration in scientometrics had been growing since 2001.

To analyze the intensity of co-authorship pattern, the whole data (816 papers) had been divided into two groups, which were single authored (271) and multi-authored (545). Compared to the result made by Dutt et al. (2003) that more than half of the papers were single authored, we found that the ratio of papers written by two or more authors had increased rapidly from 2002-2008.

Table 6 listed the TOP 10 authors according to their number of papers. Compared to Fig. 6 we could find that all the TOP 10 authors were appeared in the biggest collaboration cluster. It was interesting to note that the TOP 10 authors collaborated with each other either directly or indirectly.

Most of the TOP 20 cited references had distributed in big co-citation clusters shown in Fig. 7. ... All these four highly cited papers in the biggest cluster were focusing on the relationship between science and technology especially for the effect of science on technology. ... The second largest cluster contained 19 nodes, two of which were ranked in TOP 20. The topics were about indexes to quantify an individual's scientific research output (Hirsch, 2005). The third largest cluster included three nodes listed in TOP 20, whose topics were about author collaborations (Glanzel, 2001; Katz & Martin, 1997; Narin, Stevens, & Whitlow, 1991). There were another two clusters containing two TOP 20 nodes, and one had 10 nodes, whose topics were on co-citation networks (De Solla. Price, 1965;Small, 1973), the other had only two nodes published in Nature and Science individually with the topic of the scientific impact and wealth of nations (King, 2004; May, 1997).

The major topic were social network analysis (Wasserman & Faust, 1994), Matthew effect in science (Merton, 1968), author self-citation (Glanzel, Thijs, & Schlemmer, 2004), country research performance (Moed, 2002), evaluation indicators of publication and citation (Schubert & Braun, 1986) and the calculation of web impact factors (Ingwersen, 1998).

2014年1月15日 星期三

Chen, Y., Börner, K., & Fang, S. (2013). Evolving collaboration networks in Scientometrics in 1978–2010: a micro–macro analysis. Scientometrics, 1-20.

Chen, Y., Börner, K., & Fang, S. (2013). Evolving collaboration networks in Scientometrics in 1978–2010: a micro–macro analysis. Scientometrics, 1-20.

科學計量學(Scientometrics)利用數學、統計與資料分析方法與技術,蒐集、處理、解釋與預測學術傳播、成效、發展與動態等科技的特徵,對科技進行量化研究。就實務的技術而言,科學計量學利用書目計量的概念測量文本與資訊,並且利用科學地圖(science map)展現結果 (Börner 2010; Börner et al. 2003)。本研究利用網絡分析技術,從巨觀(國家)、中觀(機構)與微觀(作者)三種層次,探討Scientometrics期刊1978-2010年發表的2541筆論文上的合作情形。

過去對於Scientometrics有以下的相關研究,Schoepflin and Glänzel (2001) 利用1980、1989和1997三年出版的Scientometrics論文,發現科學政策(science policy)與科學社會學(the sociology of science)等主題相關的論文比率減少。Peritz and Bar-Ilan (2002) 以1990和2000年的Scientometrics論文,確認Research Policy和Social Studies of Science分別是第三和第四最常引用的期刊。Hou, Kretschmer, and Liu (2008)從Scientometrics2002到2004年論文上的作者合作網絡上發現一半以上的作者有合作的經驗,但是網絡的連結並不強而且疏鬆。Dutt, Garg, and Bali (2003)使用大量論文資料進行,研究資料期間為1978到2001,發現機構的平均論文數偏低,顯示研究的產出相當分散,而且以單一作者的論文為主,雖然多位作者的論文正蓄勢待發。

研究結果發現:
(1) 論文生產力較大的國家有美國、比利時、英國、荷蘭和西班牙。
(2) 機構與作者數隨時間增加,但是機構的平均論文數成長緩慢,近年的作者平均論文數則減少。
(3) 具有高中心性及中介性的一些機構可視為是合作網絡上的守門人(gatekeepers)。
(4) 近期較高生產力的作者取代了早期的重要作者。

Scientometrics的論文、作者和機構平均每年增加率為20%,顯示這個領域吸引愈來愈多的研究人員和機構加入。另外,Bettencourt et al. (2009) 以下面的公式指出當領域成長時,它的合作網路將會變得更為稠密。

而本研究三種層次的合作網絡,國家合作網絡的α值為 2.9533,機構與作者則分別為1.5222和1.2353。明顯的可以看出國家合作網絡相當快速地變得稠密,但由於單一作者論文的增加以及許多合作仍然是同一國家內的機構或同一機構內的作者彼此間的合作,使得機構合作網絡和作者合作網絡的α值較小。
在網絡的直徑(diameter),也就是網絡上最長的路徑方面,國家合作網絡的直徑在1989到1998年是增長的情形,但在近十年則是減短;反之,機構合作網絡和作者合作網絡到2010年仍在持續增長。
測量三種層次合作網絡的節點的連結程度(connection degree),其分布情形都符合冪次法則(power law)。節點的重要性可以從它們的程度中心性和中介中心性來推測,早期和近期的重要作者有很大的不同,在國家與機構方面的差別相當小。
比較三種層次的結果可以發現,較低層的結果會影響到上面的層次。例如有些作者的高排名不僅影響機構的排名,同時也會主導國家的排名;當具有高生產力的作者移動後,會牽動機構網絡的結構性變化。

Specifically, we would like to understand if and how collaborations at the author (micro) level impact collaboration patterns among institutions (meso) and countries (macro).

All 2,541 papers (articles, proceedings papers, and reviews) published in the international journal Scientometrics from 1978–2010 are analyzed and visualized across the different levels and the evolving collaboration networks are animated over time.

(1) USA, Belgium, and England dominated the publications in Scientometrics throughout the 33-year period, while the Netherlands and Spain were the subdominant countries;

(2) the number of institutions and authors increased over time, yet the average number of papers per institution grew slowly and the average number of papers per author decreased in recent years;

(3) a few key institutions, including Univ Sussex, KHBO, Katholieke Univ Leuven, Hungarian Acad Sci, and Leiden Univ, have a high centrality and betweenness, acting as gatekeepers in the collaboration network;

(4) early key authors (Lancaster FW, Braun T, Courtial JP, Narin F, or VanRaan AFJ) have been replaced by current prolific authors (such as Rousseau R or Moed HF).

Comparing results across the three levels reveals that results from one level might propagate to the next level, e.g., top rankings of a few key single authors can not only have a major impact on the ranking of their institution but also lead to a dominance of their country at the country level; movement of prolific authors among institutions can lead to major structural changes in the institution networks.

Scientometrics is a distinct discipline that performs quantitative studies of science and technology using mathematical, statistical, and data-analytical methods and techniques for gathering, handling, interpreting, and predicting a variety of features of the science and technology enterprise, including scholarly communication, performance, development, and dynamics.

In practice, scientometrics often requires the use of bibliometrics, the measurement of texts and information, and results might be presented as science maps (Börner 2010; Börner et al. 2003).

The study presented here uses papers that appeared in Scientometrics, the flagship journal of the field (Chen et al. 2002) publishing a major percentage of works in scientometrics as well as in the field of informetrics (Bar-Ilan 2008) over the last 33 years.

For example, Schoepflin and Glänzel (2001) used papers published in Scientometrics for the years 1980, 1989, and 1997 to identify a decrease in the percentages of both the articles related to the subjects of science policy and to the sociology of science.

Peritz and Bar-Ilan (2002) used papers published in Scientometrics for the years 1990 and 2000 and confirmed that Research Policy and Social Studies of Science are the third and fourth most frequently referenced journals in articles published in Scientometrics.

Hou et al. (2008) analyzed the structure of scientific collaboration networks in scientometrics at the micro level (individuals) by using bibliographic data of all papers published in Scientometrics from the years 2002–2004. They found that although half the authors had co-authored with each other, the network was not strongly connected and the collaborative network in the field of scientometrics was very loose.

Dutt et al. (2003) analyzed Scientometrics papers published during 1978–2001, examining the distribution of countries and themes and comparing institutions and coauthors to show that the research output is highly scattered, as indicated by the average number of papers per institution and dominated by single-authored papers; however, multi-authored papers are gaining momentum.

Chen et al. (2010) introduced a multiple-perspective co-citation analysis for characterizing and interpreting the structure and dynamics of co-citation clusters of the field of information science between 1996 and 2008. He showed that the multiple-perspective method increases the interpretability and accountability of both author-citation analysis (ACA) and document- citation analysis (DCA) networks.

Wagner and Leydesdorff (2005) applied network analysis to map the growth of international co-authorships, and they found that international co-authorships can be explained based on the organizing principle of preferential attachment, although the attachment mechanism deviates from an ideal power-law.

Samoylenko et al. (2006) visualized the scientific world and its evolution by constructing minimum spanning trees (MSTs) and a two-dimensional map of scientific journals using the Science Citation Index from the Web of Science database for 1994–2001 and showed a linear structure of the scientific world with three major domains: physical sciences, life sciences, and medical sciences.

Perc (2010) studied the evolution of Slovenia’s scientist collaboration network from 1960 to 2010 with a yearly resolution and showed the network had a ‘‘small world’’ pattern and its growth was governed by near-linear preferential attachment. This paper will advance the existing works by studying the evolution of scientometrics at three different network levels.

Figure 1 shows the growth (annual and cumulative) of the number of papers, countries (or regions), institutions and authors from 1978 to 2010. By counting the annual numbers in each figure, we obtain average annual growth rates, which are 20.4 % (papers), 9.4 % (countries), 19.6 % (institutions), and 20.1 % (authors).

As Bettencourt et al. (2009) pointed out, when fields grow, their collaboration networks densify—i.e., the average number of edges per node increases over time. They found that the relation between the number of nodes and edges followed a simple scaling law with scaling exponent (α > 1):



Figure 2 shows that the scaling exponent a equals 2.9533 at the macro-country, 1.5222 at the meso-institution, and 1.2353 at the micro-author levels. It has the highest value for countries—i.e., the country collaboration networks densify rather quickly, which is also due to the fact that this is the network with the fewest nodes. However, a large number of within-country or within-institution collaborations or an increase in single-authored papers would also result in smaller α values.

The diameter of a collaboration network has major implications for information diffusion—the shorter a pathway of coauthor linkages that connects an author pair, the more likely knowledge diffuses.

Over the 33 years, the country collaboration network diameter grew from 1989 to 1998 (there were no edges before 1989), achieves the highest value in 1998, and decreases in the last 10 years. This might be due to the rather limited number of countries that perform scientometrics research.

The diameters of the institution and author collaboration networks increase continually and both reach a diameter d = 15 in 2010.

A closer look at the density of the three networks (the ratio of the number of actual edges to all possible edges in a fully connected graph with the same number of nodes) shows that both the meso and micro networks’ densities decrease over time while the macro network, which experienced a topological transition from large to decreasing diameter, shows an increase in density.

In an attempt to understand the structure of the 1978–2010 networks, the degree for each node in the network was determined and the node degree distribution p(k) plotted in Fig. 3. ... All three networks exhibit power law degree distributions.

To understand which countries, institutions, and authors play key roles in the three networks, the degree centrality (the number of links a node has) and betweenness centrality (nodes that have a high probability to occur on a randomly chosen shortest path between two randomly chosen nodes have a high betweenness) (Freeman 1977) values for each node were calculated. The resulting TOP-5 countries, TOP-10 institutions, and TOP-10 authors calculated for every 6 years (cumulatively from 1978) are listed in Tables 1 and 2.

In addition, the last table column shows the TOP-10 countries, institutions, and authors if only 2001–2010 data is considered. While the differences are minimal for countries and institutions, the list of TOP-10 authors changes considerably if only recent works are considered.

Figure 4 shows that, by the end of 2010, Belgium, USA, England, Germany, the Netherlands, China, and France are central network nodes with a large number of papers. These six countries not only link to each other but also to outside countries—e.g., Belgium and Germany have strong links to Hungary, and Belgium and England have strong links to Finland.

When analyzing the evolving institution collaboration networks, it becomes clear that a few key institutions manage to stay in the TOP-10 list—among them are the Univ Sussex, KHBO, Katholieke Univ Leuven, Hungarian Acad Sci, and Leiden Univ.

During the evolution of the co-author networks, early authors are replaced by current authors. Most TOP-10 authors from 1980 and 1986 are missing in the later years. Key authors listed in the TOP-10 lists around 1986 decline in ranking or are replaced by other authors.

One might assume that rankings on the author (micro) level impact the ranking of institution (meso) and country (macro) levels. While author rankings impact institution rankings; institution rankings are less predictive of country rankings, as exemplified below.

As can be seen in Table 3, USA ranks first in the number of institutions and the number of papers over the 33 year time span. However, the average number of papers per institution was low for the USA, especially when compared with Belgium, Netherlands, and Hungary. ... Similarly, while no single author in the USA appears in the TOP-10 lists, the number of all authors combined and the number of their papers results in a high country ranking.

Can one single author impact the ranking of an entire institution or country? The answer is yes. ... The 155 papers of the Hungarian Academy of Sciences were co-authored with 30 institutions, 22 of which were contributed by papers authored by Glänzel W. As for the 93 papers by the Katholieke Universiteit Leuven, 13 of 51 institution links were added by Glänzel W.

Over the 33 years, the number of countries grew steadily with a linear growth feature with USA, Belgium and England leading in terms of centrality and betweenness. ... As their share increases, they have a stronger impact on the evolution of scientometrics. Over time, more and more collaboration links are generated and the average node degree and network density increase as well (see Table 4).

It is important to point out that some top-ranking countries have a small number of top-ranking institutions (e.g., Katholieke Univ Leuven in Belgium) while other countries (USA) have a large number of contributing institutions.

Similarity, some top-ranking institutions have one or two top-ranking authors, e.g., Glänzel W and Rousseau R

That is, single authors can not only have a major impact on the ranking of their institution but also of their country.

At the same time, the growth rate of institutions, authors and papers for each year were similar about 20 %. It suggested that this field had been attracting more and more institutions and authors to join the field of scientometrics.

The co-author network analysis showed that many new authors joined the field of scientometrics, especially in the recent 8 years. The diameter, average degree, and density of the network show the same trends as those calculated for institutions.

While co-author networks experience the departure of senior and the arrival of young researchers, the institution and country networks seem to have a comparatively stable structure of key nodes.

2013年12月2日 星期一

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

Lu, K., & Wolfram, D. (2012). Measuring author research relatedness: A comparison of word‐based, topic‐based, and author cocitation approaches. Journal of the American Society for Information Science and Technology, 63(10), 1973-1986.

科學映射圖(scientific mapping)能夠科學結構(scientific structure)視覺化,幫助使用者確認科學主題(scientific themes)並從而發現新知識的有用工具之一。過去的研究曾經使用過作者、文章與等映射單位。在計算映射單位之間的關連,Börner, Chen, and Boyack (2005) 將關連性的測量方法(relatedness measures)分為引用連結(citation linkages)與共現相似性(co-occurrence similarities)等兩大類,而本研究則將目前常用來評估作者間的關連分為直接引用(direct citation)、共被引分析(cocitation analysis)、合著分析(co-authorship analysis)、書目耦合分析(bibliographic coupling analysis)以及共詞分析(co-word analysis)等五種方法。也有研究以發展出整合文字內容與連結的測量方法來計算期刊(Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005)與文章(Liu et al., 2010)間的關連。本研究建議兩種以詞語為基礎並利用向量空間模式(vector space modeling)的方法和另一種基於LDA(latent Dirichlet allocation)的主題模型方法來測量作者之間的關連。本研究將第一種方法稱為靜態(static)的特徵,以每位作者曾寫過的論文內容為基礎產生代表這位作者的特徵向量,也就是代表這位作者的特徵向量是所有他寫過的論文的特徵向量總和,任何兩位作者之間的關連是對應於他們的作者特徵向量之間夾角的餘弦值(cosine value)。第二種方法則是動態(dynamic)的特徵,如果兩位作者之間沒有合著的論文,他們之間的關連仍然是他們的作者特徵向量之間夾角的餘弦值,但如果他們曾經合著過,在計算他們之間的關連時,先將他們合著論文的特徵向量排除在他們的作者特徵向量之外,在進行餘弦值計算,所以在計算每位作者和其他作者之間關連時所使用的作者特徵向量可能是變動的,因此稱為動態。基礎的主題模型假設每一個論文都是主題的混合(mixture),而每一個主題則都是詞語的混合。對於每一個論文,它的主題混合由一個已知參數α的Dirichlet分布所產生;每一個主題的詞語混合則由另一個已知參數β的Dirichlet分布所產生。在產生論文d前先根據Dirichlet分布Dir(α)取樣產生它的主題混合θd,然後再產生這個論文裡的每一個詞語,每一個詞語的產生是根據從主題混合θd中取樣得到的主題z以及其相對應的詞語混合ϕk所產生。本研究採用Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010)將作者資訊加入而擴充的LDA模型-- 作者-主題模型(author-topic model),這個模型假定每個作者是由一個已知參數α的Dirichlet分布所產生的主題混合。假設一個論文的作者群為ad ,在產生這個論文的每一個詞語時,首先從ad 中隨機抽取一個作者x以及他的主題混合θx,然後其主題z便由θx取樣產生。本研究利用Gibbs取樣(Gibbs sampling, Griffiths & Steyvers, 2004)進行作者-主題模型推論,產生包含每一個主題在詞語上的分布情形以及對每一位作者產生他在各主題上的分布情形等結果。因此利用作者-主題模型可以根據他們在主題分布的相似度測量他們的關連。

本研究的資料範圍為2000到2010年出版的圖書資訊學相關的八種主要期刊的 5227筆書目紀錄,從其中的 6282位不同的作者內選取50位最多產的作者。利用靜態特徵、動態特徵、主題模型和共被引分析等四種方法測量多產作者之間的關連並利用MDS (multidimensional scaling)和階層式叢集分析(hierarchical cluster analysis)進行視覺化。本研究在利用主題模型測量作者之間的關連時使用以下的參數,α設為50/K,其中的K是主題的數量,本研究設為20,β設為0.01,Gibbs取樣的迭代(iteration)次數設為1000次。針對每一對作者的四種關連測量方法所得到的值進行相關分析(correlation analysis),結果發現靜態特徵與動態特徵之間有最高的相關值,主題模型和其他兩種以內容為基礎的測量方法的相關值也較共被引方法來得高。四種測量方法皆可以發現LIS領域的兩大主軸:一個主軸是資訊檢索(information retrieval)與網路研究(web studies),另一則是科學評鑑(scientific evaluation)的測量指標(metrics)研究,LDA模型則在階層式叢集分析上有最連貫的結果。另外,以內容為基礎的方法比以引用為基礎的方法更容易解釋產生的結果。

In this study we present static and dynamic word-based approaches using vector space modeling, as well as a topic-based approach based on latent Dirichlet allocation for mapping author research relatedness.

Outcomes for the two word-based approaches and a topic-based approach for 50 prolific authors in library and information science are compared with more traditional author cocitation analysis using multidimensional scaling and hierarchical cluster analysis.

Science mapping is one of the most useful tools to visualize scientific structure. It helps to identify scientific themes, and discover new knowledge.

The unit of interest for mapping may include authors, articles, and journals.

To date, five approaches have been used to measure the relatedness between authors, where the nature of the relationship studied is based on the data used: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling analysis, and co-word analysis.

Recently, more sophisticated hybrid methods (i.e., using textual content and citations) have been applied to the mapping of articles (Ahlgren & Colliander, 2009; Boyack & Klavans, 2010; Cao &Gao, 2005) and journals (Liu et al., 2010).

As an initial investigation of these topics, our focus will be on authors whose publications appear in the highest impact library and information science journals.

In reviewing visualization studies for knowledge domains, Börner, Chen, and Boyack (2005) categorized relatedness measures into two broad categories: citation linkages and co-occurrence similarities.Within the relatedness measures, five basic approaches were identified: direct citation, cocitation analysis, co-authorship analysis, bibliographic coupling, and co-word analysis.

Direct citation accounts for the relatedness between a citing work and a cited work based on citing behavior. ... Shibata, Kajikawa, Takeda, and Matsushima (2008) explored citation networks for two research domains and divided the networks into clusters in order to identify research fronts. Direct citation has not attracted wide attention. One possible reason may be its requirement for a very long time window to obtain a sufficient linking signal for clustering (Boyack & Klavans, 2010).

The idea that two articles that share the same references are related, referred to as bibliographic coupling, was outlined by Kessler (1963). The more references two articles have in common, the more closely related they are thought to be. Note that this list is static over time because references within articles do not change. With the interrelation of this link, scientific products can be ordered into groups. Weinberg (1974) reviewed the theory and practical applications of bibliographic coupling and granted the usefulness of the method. More recently, Zhao and Strotmann (2008) aggregated bibliographic coupling at an author’s oeuvre (body of work) level, which they called author bibliographic-coupling analysis (ABCA). They found ABCA can provide an effective picture of current active research in a field.

Cocitation analysis, introduced by Small (1973), is probably the most influential approach for assessing relatedness measures. If two articles are cited by the same third article, these two articles are co-cited. The assumption is that the appearance of two articles in the same reference list indicates a semantic association between the articles. Unlike traditional bibliographic coupling, cocitation is a dynamic relationship based on the citing authors. New citing authors can change the cocitation relationship. This feature is important because science is developing continuously. Relationships among scientific units being studied should be able to incorporate this dynamic change.

White and Griffith (1981) first applied cocitation techniques to authors, called author cocitation analysis or ACA. The essential transformation is to consider “Author” as a body of writings by a person (i.e., an oeuvre). So the cocitation of authors applies to any work by any author being co-cited with any work by another author.

Since then, a number of studies have been conducted using variations of the ACA method, including normalization (Ahlgren, Jarneving,&Rousseau, 2003; Leydesdorff&Vaughan, 2006; White, 2003; van Eck & Waltman, 2009), author counts (Zhao & Strotmann, 2011), and last-author ACA (Zhao & Strotmann, 2010).

One disadvantage of cocitation analysis is the lack of cognitive interpretation of the relatedness of the co-cited units. Without enough domain knowledge, one can hardly interpret the cocitation map.

Leydesdorff (1987) argued that cocitation maps only partially represent the structure of science.

A co-authorship relationship is established when authors co-publish a paper. Glänzel (2001) studied international co-authorship links to reveal the structures in international collaborations. Liu, Bollen, Nelson, and Van de Sompel (2005) constructed a network with co-authorship relations in the field of digital libraries. Ding (2011b) studied scientific collaborations and citation patterns of researchers and combined the results with a topic model approach to examine collaborations among researchers who share similar and different research interests.

It is this feature of co-authorship that makes co-authorship analysis more revealing of a social network rather than a scientific structure.

Co-word analysis collects evidence of relatedness from co-occurring keywords from different articles. Compared with the approaches introduced earlier, co-word analysis directly uses actual contents to measure relatedness, whereas the others find indirect evidence through citation and co-author relations. An obvious advantage of co-word analysis is that relatedness can be interpreted directly according to document contents.

Coulter, Monarch, and Konda (1998) mapped the discipline of software engineering with co-word analysis. Indexing terms from the ACM Computing Classification System were used as the unit of analysis. Ding, Chowdhury, and Foo (2001) conducted a co-word analysis on a sample of 2,012 articles from the Web of Science (WoS) to reveal themes of information retrieval research.

Leydesdorff (1997) noted that the meaning of words change from position to position and from one text to another. He also suggested this change will destabilize the science map produced by co-word analysis.

Another disadvantage of using indexer-assigned keywords as the source for co-word analysis is the “indexer effect” (Law & Whittaker, 1992), which creates bias through factors such as the artificiality of an indexing language, delays in changes to the indexing language to reflect the current state of a discipline, and subjectivity in the assignment of index terms.

In the vector space, a number of documents constitute a document space. The centroid of the document space is a summarization of the characteristics of the space. It represents the average vector for a group of documents.

Each author will be viewed as a document space consisting of the articles he/she has written. This space is a subspace of the collection space, named the author space. The centroid of the author space will be used to represent the author. The relatedness between authors will be measured through the similarity between the centroids of their author spaces.

The topic model is an improvement over the basic vector space model in terms of relieving the independence assumption and capturing the term associations. Instead of assuming independence among terms, the topic model assumes exchangeability among terms in documents, which is a much looser assumption.

Early works on the topic model include latent semantic indexing (LSI) by Deerwester et al. (1990) and the probabilistic LSI (pLSI) by Hofmann (1999). LDA is a more recent technique proposed by Blei, Ng, and Jordan (2003). It has an advantage over LSI in explicitly modeling the latent topics, and over pLSI in solving the overfitting problem (i.e., a model with too many parameters).

The LDA model treats a document as a mixture of topics and a topic as a mixture of terms. Each document (i.e., a mixture of topics θ) is generated from a latent Dirichlet distribution with a prior of α, and each topic (i.e., a mixture of terms ϕk) is generated from a Dirichlet distribution with a prior of β. The generation process entails, first, sampling a document θd from Dir(α). At each position of a word in a document, a topic z is selected according to θd, and a word w is selected according to z and ϕk.



Rosen-Zvi, Chemudugunta, Griffiths, Smyth, and Steyvers (2010) extended the original LDA model to include authors and proposed the author-topic model (Figure 2). This model includes authorship information in the generative process. Each document has a number of authors ad. Each author is considered as a distribution of topics drawn from a Dirichlet distribution with a prior of α. For each word in a document, an author x is randomly drawn from ad and the topic distribution associated with this author is θx. Then a topic z is selected the same way as in a LDA model to generate the observed word w.



The advantage of this author-topic model is that it adds authorship information to the model, so that the topics are learned and assigned to documents accordingly. In the output of this model, each author is a distribution of different topics; each topic is a distribution of terms. As the purpose of the current study is to measure the relatedness of authors, the author-topic model will be appropriate to produce author similarities based on their topics.

Gibbs sampling (Griffiths & Steyvers, 2004) is used to estimate the parameters in the model.

Table 1 lists the eight journals selected for inclusion in the study. ... Bibliographic records for documents published in these journals between 2000 and 2010 were downloaded. Records downloaded were further limited to three document types: articles, proceedings papers, and reviews. ... In total, 5,227 records were downloaded from WoS. The raw WoS records were processed, and only three fields were kept: the article title (i.e., “TI” field), the Keywords Plus (i.e., “ID” field), and the abstract (i.e., “AB” field). The records then were indexed with the widely used Lemur information retrieval toolkit (http://www.lemurproject.org/). Stop words were removed and stemming was applied.



From the 5,227 records downloaded, we were able to identify 6,282 different author names using string matching. Because it is impractical to map all of the authors in our collection, we selected the 50 most prolific authors according to the WoS “analyze results” function. ... We selected the most prolific authors because the more an author writes, the better the algorithm used “understands” her/his interests, and thus the more accurate our assessment will be.

For each author in our author list we then generated an author space consisting of all the articles he/she wrote. TF*IDF term weighting was employed to assign term significance in the space. Terms that were single characters or only consisted of digits (e.g., “2001”) were filtered out. We believe that these terms add noise to the space rather than meaning. The relatedness between authors is measured through the cosine between the centroids of the author spaces.

One could argue that this creates a biased assessment of the strength of the relationship because there is an exact match for the text of the co-authored publications that creates a stronger bond than for two authors who have published in a common area but did not collaborate. On the other hand, the simple fact that the collaboration has resulted in one or more co-authored documents should be acknowledged as a strong tie between the authors.

In a static space, each author has her/his own space that consists of her/his articles. This space does not change when measuring author relatedness. ... The relatedness of authors will include the similarity arising from the strength of the co-authorships.

Conversely, in the dynamic author space, the author spaces depend on a pair of authors. Co-authored articles by the pair of authors are excluded. In this case, each author may have a different author space when measured with different authors.

The vector space model provides a number of readily available measures of relatedness. The most popular is the cosine measure, which measures the cosine of the angle formed by two vectors in the space. It basically measures the term weight distribution between two vectors. The more similar the distribution is, the higher the cosine value is expected to be.

Gibbs sampling (Griffiths & Steyvers, 2004) was used to estimate the parameters in the author-topic model. We set the number of iterations to 1,000. The hyperparameter α was set to 50/K where K is the number of topics and hyper β is set to 0.01.We tested different K, or number of topics, values and decided to report the results from K = 20 because it produced the most reasonable outcome by our judgment.

The topic model toolbox was employed to perform the learning process (http://psiexp.ss.uci.edu/research/programs_data/toolbox.htm).

An author-topic LDA model (Rosen-Zvi et al., 2010) was trained on our collection and a pair-wise cosine similarity measure comparison of the 50 authors was conducted, resulting in a symmetric matrix of similarity values based on the LDA modeling. Similarity matrices were also calculated for both the static and dynamic author spaces. Multidimensional scaling was used to visualize the relationships among the authors. ... Because the data represent a type of similarity measure, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006). To provide additional insights into the grouping of the authors, hierarchical cluster analysis (complete linkage method) was used in SPSS to superimpose groups of authors on the MDS maps to provide an additional means to assess the coherence in the resulting proximities between authors.

After tokenization of the field contents, 916,383 tokens, or individual words, were identified; the number of unique tokens, or distinct words, was 12,537. The average document length was 175.32 tokens.

An examination of the pair-wise correlation of these author relatedness measures reveals significant and moderate level correlations between the word-based, topic-based, and author cocitation measures (Table 5). It is not surprising that the static author map has a high correlation with the dynamic author map (Kendall’s tau b = 0.971). Similarly, the correlations among the three content-based approaches are generally higher than their correlations with the cocitation approach. This provides preliminary evidence that they measure different types of relationships.

In all cases, the largest singular group consists of authors who work with different aspects of metrics-based studies, which is labeled as “Informetrics” in general in the two word-based maps and “Scientific impact evaluation” in the other two maps. This labeling indicates that the metrics-related topics have been a frequently investigated theme by the prolific authors in the selected journals during the first decade of the 21st century.

It is also noteworthy that the topic groupings of each of the maps largely aligns along the horizontal or vertical axis, with one side representing information retrieval (system and behavior) and web studies, with the other side corresponding to metrics-based or scientific evaluation studies.

As is shown from the maps, the static map (Figure 3) and dynamic map (Figure 4) are generally consistent in terms of the location of the authors, which indicates that the exclusion of similarities resulting from collaborations does not affect the overall layout. However, drastic changes may happen to individuals who have collaborated frequently with another author.

At the four-cluster agglomeration, the LDA map (Figure 6) provides the most coherent representation of the author map in relation to the generated clusters. At the two-cluster agglomeration, the clusters are neatly divided along the vertical axis, with metrics-related research represented on the left, and web and information retrieval-related themes on the right. Although the group membership of some individuals is still debatable, such as “Ingwersen_P” in the “Scientific impact evaluation” group given that he has also published in information retrieval and webometrics, the overall layout of the LDA map does provide semantically meaningful relationships.

Of the five author relatedness methods discussed earlier, only co-authorship provides a direct connection between authors.

Cocitations are contributed by third parties.

Direct citations reflect an author’s assessment of relatedness to a cited author or work but are still based on perception or the subjectivity inherent in citer motivation (Bornmann & Daniel, 2008).

This is also the case for bibliographic coupling, where the strength of the relationship is assessed by the overlap of references selected by two authors.

Co-word or topic-based studies can be argued to be the least influenced by citing behavior because they rely solely on the words developed by the authors themselves.

The newly proposed content-based approaches overcome several limitations of the more traditional cocitation approach.

In addition to avoiding citer subjectivity inherent in citation-based data, the links between authors will be more interpretable compared with the cocitation maps. The top terms/topics will be identifiable to help interpret the links between authors.

The content-based methods do not require an author to be cited in order to be included in the map. As long as the author has some publication record, her/his relatedness with other authors can be identified. This provides the opportunity for researchers who have not been widely cited to be included in the author map.

Furthermore, cocitation analysis outcomes may be affected by limited numbers of citations that do not reflect the true strength of the relationship between authors. This can be seen when comparing the cocitation outcomes with the topic-based outcomes, where several authors with low citation counts, and therefore low cocitation counts, end up at the periphery of the map. For the LDA outcome, these authors are more centrally situated among authors with similar topic areas.

The word-based and topic-based methods can be considered an extension of co-word analysis, where words are used to determine the relatedness of authors.

In Healey, Rothman, and Hoch (1986), a paradox is introduced: if a map represents a field that is already known to experts, then it is useless because it does not reveal anything new; if the map deviates from the expectation of the experts, then its outcome is questionable.

This initial investigation, which compares prolific authors from LIS, demonstrates: (1) the potential for more topically meaningful outcomes from the new methods when compared to more traditional cocitation analysis; (2) the topicbased method using LDA for the data used in this study produces more distinctive clusters and reasonable results than the two word-based approaches.