顯示具有 procedure 標籤的文章。 顯示所有文章
顯示具有 procedure 標籤的文章。 顯示所有文章

2015年4月15日 星期三

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V. (2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V.(2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

由於一般認為將領域之間的關係表示為圖形,通過考慮這些關係的可能性能夠提供許多資訊,不論對新進人員或專家皆有助於理解與分析,因此對這方面方法與工具的需求逐漸提高。過去的研究大多以期刊為分析單位,產生所有科學研究領域的科學映射圖。例如Leydesdorff (2004a, 2004b)使用雙重連結成分(biconnected components)的圖形分析演算法,將JCR 2001的科學研究進行分類。Boyack, Klavans, and Börner (2005)則應用了8種不同的期刊相似性測量7121種SCI和SSCI期刊,並採用VxOrd產生科學映射圖。Samoylenko, Chao, Liu, and Chen (2006)建構科學期刊的最小生成樹(minimum spanning trees),他們使用的資料是SCI 1994到2001的資料。本研究提出一個將ISI (Institute of Scientific Information)類別繪製成科學映射圖的方法,這個方法利用根據類別間的共被引資訊建構類別間的連結,以尋徑網路(PathfinderNetwork)縮減不重要的連結,然後以Kamada-Kawai方法決定節點在圖上的布局(layout),最後利用因素分析(factor analysis)進行結構確認。本研究和先前的研究都是針對類別利用共被引資訊呈現科學映射圖。以類別為分析單位在代表上足夠明確,並且比起較小的單位,這種方式對非專家使用者(nonexpert user)較具有資訊且使用者友善。Moya-Anegón et al. (2004)針對西班牙科學研究領域的視覺化,Moya-Anegón et al. (2005)則進一步利用科學映射圖比較英國、法國和西班牙三個國家的科學研究領域。本研究依循Börner, Chen, and Boyack (2003)提出的知識領域映射流程。使用的資料為7585種ISI期刊,ISI的類別共有219個,但扣除多學科科學後(Multidisciplinary Sciences),採用的類別共218個。利用共被引計算期刊相似性的方式為

Cc(ij)為期刊i和期刊j共被引次數,c(i)和c(j)則分別是期刊i和期刊j被引用次數。然後以尋徑網路和Kamada-Kawai方法繪製網路圖,經過尋徑網路處理後,有較多連結的節點具有較重要的地位。而尋徑網路是一種以型態為主的方法,與以群集為主的因素分析彼此間可以互補,因素分析可以識別、界定與定名科學映射圖上呈現的主題區域,而尋徑網路則負責讓使主題區域更加明顯,將類別分組成束,並顯示連接不同顯著類別的路徑,以及總體的型態結構。。最後總計共分析出35個因素,通過陡坡考驗(scree test)則有16個。科學映射圖上的類別可以分為三個群集:醫學與地球科學、基礎與實驗科學以及社會科學。

This study proposes a new methodology that allows for the generation of scientograms of major scientific domains, constructed on the basis of cocitation of Institute of Scientific Information categories, and pruned using PathfinderNetwork, with a layout determined by algorithms of the spring-embedder type (Kamada–Kawai), then corroborated structurally by factor analysis.

We present the complete scientogram of the world for the Year 2002.

This need arises from the general conviction that an image or graphic representation of a domain favors and facilitates its comprehension and analysis, regardless of who is on the receiving end of the depiction and whether a newcomer or an expert.

Science maps can be very useful for navigating around in scientific literature and for the representation of its spatial relations (Garfield, 1986). They are optimal means of representing the spatial distribution of the areas of research while also offering additional information through the possibility of contemplating these relationships (Small & Garfield, 1985).

From a general viewpoint, science maps reflect the relationships between and among disciplines; but the positioning of their tags clues us into semantic connections while also serving as an index to comprehend why certain nodes or fields are connected with others.

Moreover, these large-scale maps of science show which special fields are most productively involved in research—providing a glimpse of changes in the panorama—and which particular individuals, publications, institutions, regions, or countries are the most prominent ones (Garfield, 1994).

It is a tool in that it allows the generation of maps, and a method in that it facilitates the analysis of domains, by showing the structure and relations of the inherent elements represented. In a nutshell, scientography is a holistic tool for expressing the discourse of the scientific community it aspires to represent, reflecting the intellectual consensus of researchers on the basis of their own citations of scientific literature.

In Moya-Anegón et al. (2004), we ventured forth with a historic evolution of scientific maps from their origin to the present, and proposed ISI-JCR category cocitation for the representation of major scientific domains. Its utility was demonstrated by a visualization of the scientific domain of geographical Spain for the Year 2000.

Since then, other works related with the visualization of great scientific domains have appeared; however, all use journals as the unit of analysis, with the exception of a study based on the cocitation of categories (Moya-Anegón et al., 2005), comparatively focusing on three geographic domains (England, France, and Spain).

In contrast, Leydesdorff (2004a, 2004b) classified world science using the graph-analytical algorithm of biconnected components in combination with JCR 2001.

Boyack, Klavans, and Börner (2005) applied eight alternative measures of journal similarity to a dataset of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes, to show the first global map of science using the force-directed graph layout tool VxOrd.

Samoylenko Chao, Liu, and Chen (2006) proposed an approach through the construction of minimum spanning trees of scientific journals, using the Science Citation Index from 1994 to 2001.

In processing and depicting the scientific structure of great domains, we further developed a methodology that follows the flow of knowledge domains and their mapping as proposed by Börner, Chen, and Boyack (2003).

Because ISI assigns each journal to one or more subject categories, to designate a subject matter (i.e., ISI category) for each document, we also downloaded the Journal Citation Report (JCR; Thomson Corporation, 2005a), in both its Science and Social Sciences editions, for 2002.

The downloaded records were exported to a relational database that reflects the structured information of the documents. This new repository contained nearly 1 million (N = 901,493) source documents: articles, biographical items, book reviews, corrections, editorial materials, letters, meeting abstracts, news items, and reviews that had been published in 7,585 ISI journals (N = 5,876 + 1,709). These were classified in a total of 219 categories, altogether citing 25,682,754 published documents.

As informational units, they are, in themselves, sufficiently explicit to be used in the representation of all disciplines that make up science in general. These categories, in combination with the adequate techniques for the reduction of space and the representation of the information to construct scientograms of science or of major scientific domains, prove much more informative and user friendly for quick comprehension and handling by nonexpert users than those obtained by the cocitation of smaller units of cocitation.

For these reasons, we used the 219 categories of the JCR 2002 as units of measure, with the exception of “Multidisciplinary Sciences.” ... The maximum number of categories with which we worked, then, was 218.

In light of our previous experience (Moya-Anegón et al., 2004, 2005), we use cocitation as the similarity measure to quantify the relationship existing between each one of the JCR categories.

Therefore, after a number of trials, we arrived at the conclusion that using tools of Network Analysis, the best visualizations are those obtained through raw data cocitation as the unit of measure. Yet, it also was necessary to reduce the number of coincident cocitations to enhance pruning algorithm yield. Therefore, to those raw data values we added the standardized cocitation value. In this way, we could work with raw data cocitation while also differentiating the similarity values between categories with equal cocitation frequencies. The key was a simple modification of the equation for the standardization of the degree of citation proposed by Salton and Bergmark:




where CM is cocitation measure, Cc is cocitation frequency, c is citation, and i and j are categories.

Over the history of the visualization of scientific information, very different techniques have been used to reduce n-dimensional space. Either alone or in conjunction with others, the most common are multidimensional scaling, clustering, factor analysis, self-organizing maps, and PathfinderNetworks (PFNET).

In our opinion, PFNET with pruning parameters r = ∞, and q = n − 1 is the prime option for eliminating less significant relationships while preserving and highlighting the most essential ones, and capturing the underlying intellectual structure in a economical way.

Although PFNET has been used in the fields of Bibliometrics, Informetrics, and Scientometrics since 1990 (Fowler & Dearhold, 1990), its introduction in citation was due to the hand of Chen (1998, 1999), who introduced a new form of organizing, visualizing, and accessing information. The end effect is the pruning of all paths except those with the single highest (or tied highest) cocitation counts between categories (White, 2001).

The spring embedder type is most widely used in the area of documentation, and specifically in domain visualization. Spring embedders begin by assigning coordinates to the nodes in such a way that the final graph will be pleasing to the eye (Eades, 1984). Two major extensions to the algorithm proposed by Eades (1984) have been developed by Kamada and Kawai (1989) and Fruchterman and Reingold (1991).

While Brandenburg, Himsolt, and Rohrer (1995) did not detect any single predominating algorithm, most of the scientific community goes with the Kamada–Kawai algorithm. The reasons upheld are its behavior in the case of local minima, its capacity to minimize differences with respect to theoretical distances in the entire graph, good computation times, and the fact that it subsumes multidimensional scaling when the technique of Kruskal and Wish (1978) is applied.

We can effortlessly see which are the most important nodes in terms of the number of their connections and, in turn, which points act as intermediaries with other lines, as hubs or forking points.

Whereas factor analysis is a clustering-oriented procedure, PFNET is topology oriented. Yet, they are extremely valuable as complements in the detection of the structure of a scientific domain.

Thus, factor analysis is responsible for identifying, delimiting, and denominating the great thematic areas reflected in the scientogram.

Meanwhile, PFNET is in charge of making the subject areas more visible, grouping their categories into bunches, and showing the paths that connect the different prominent categories, and finally, the overall topology of the domain.

Factor analysis identifies 35 factors in the cocitation matrix of 218 × 218 categories of world science 2002. Through the scree test we extracted 16, which we tagged using the previously explained method; these accumulate 70.2% of the variance (Table 1)

The number of categories included in at least one factor is 195. Twenty-three were not included in any factor (Table 2), and 25 belonged to two factors simultaneously (Table 5).

That is, a category or thematic area occupying a central position in the scientogram will have a more general or universal nature in the domain as a consequence of the number of sources it shares with the rest, contributing more to scientific development than those with a less central position.

The more peripheral the situation of a category or subject area, the more exclusive its nature, and the fewer the sources it will appear to share with other categories; accordingly, the lesser its contribution to the development of knowledge through scientific publications.

An intermediary position favors the interconnection of other categories or thematic areas. 

This broad interpretation of our scientograms not only explains the patterns of cocitation that characterize a domain but also foments an intuitive way for specialists and nonexperts to arrive at a practical explanation of the workings of PFNET (Chen & Carr, 1999).

From a macrostructural point of view, we can distinguish three major zones.

In the center is what we could call Medical and Earth Sciences, consisting of Biomedicine, Psychology, Etiology, Animal Biology & Ecology, Health Care & Service, Orthopedics, Earth & Space Science, and Agriculture & Soil Sciences.

To the right, we can see some other basic and experimental sciences: Materials Sciences & Physics, Applied; Engineering; Computer Science & Telecommunications; Nuclear Physics & Particles & Fields; and Chemistry.

To the left is the neighborhood of the social sciences, with Applied Mathematics, Business, Law, and Economy, and Humanities.

On one hand, it offers domain analysts the possibility of seeing the most essential connections between categories of given domain.

On the other hand, it allows us to see how these categories are grouped in major thematic areas, and how they are interrelated in a logical order of explicit sequences.

2014年2月28日 星期五

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of American Society for Information Science and Technology, 57(3), 359-377.

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature.  Journal of American Society for Information Science and Technology, 57(3), 359-377.

information visualization

本研究提出一個整合研究專業(specialty)的研究前沿(research front)以及其引用的知識基礎(intellectual base)的視覺化介面。本論文定義研究前沿為研究專業上一組急遽出現的概念(concepts)與研究議題(research issues);研究前沿的知識基礎則是包含這些概念與研究議題的論文引用或者共同被引用的論文。在針對某一個專業進行其研究前沿與知識基礎進行視覺化時,首先蒐集專業相關的論文,從這些論文抽取代表研究前沿的詞語,並以論文所引用或共被引的論文做為專業的知識基礎,建立分別代表研究前沿的詞語和知識基礎的論文的二方網路(bipartite networks)以同時呈現研究前沿的相關概念與研究議題以及知識基礎的論文。在建立起來的網路上透過詞語和論文形成的叢集可以發現重要的研究前沿和知識基礎,藉由詞語呈現叢集的概念與研究議題更能有效地表達研究前沿的意涵,並且如果加上論文的發表時間來分析,可以從急遽出現在較多論文的相關詞語找出發展中的研究前沿。此外,對於網路進行中介中心性(centrality of betweenness)分析可以發現研究前沿間具有樞紐地位的論文,並且透過Pathfinder演算法可以發現論文間的主要關連。
A specialty is conceptualized and visualized as a time-variant duality between two fundamental concepts in information science: research fronts and intellectual bases.
A research front is defined as an emergent and transient grouping of concepts and underlying research issues.
The intellectual base of a research front is its citation and co-citation footprint in scientific literature— an evolving network of scientific publications cited by
research-front concepts.
The concept of a research front was originally introduced by Price (1965) to characterize the transient nature of a research field. Price observed what he called the immediacy factor: There seems to be a tendency for scientists to cite the most recently published articles. In a given field, a research front refers to the body of articles that scientists actively cite.
A specialty can be conceptualized as a time-variant mapping from its research front to its intellectual base.
Typical questions regarding a research front may include:
How did it get started? What is the state of the art? What are the critical paths in its evolution?
To address such questions, we need to detect and analyze emerging trends and abrupt changes associated with a research front over time. We also need to identify the focus of a research front at a particular time in the context of its intellectual base, to reveal significant intellectual turning points as a research front evolves, and to discover the interconnections between different research fronts.
Braam, Moed, and Raan (1991) defined a specialty as “focused attention by a number of scientific researchers to a set of related research problems and concepts” (p. 252). They studied the continuity and stability of a specialty in terms of the similarity between co-citation clusters across consecutive years. The similarity between two co-citation clusters is determined by comparing aggregated word profiles of the clusters.
In part, this is because we define a research front differently to emphasize emerging trends and abrupt changes as the defining features of a research front. A research front is the domain of a time-variant mapping, and its intellectual base is the co-domain of the mapping.
Griffith et al. (1974) found that between-cluster co-citation links tend to be weaker than within-cluster co-citation links. ... To understand how specialties and different thematic trends interact with each other, it is essential to study the nature of long-range, between-cluster links and understand why articles in different specialties were connected.
Labeling clusters is concerned with the clarity and interpretability of co-citation clusters. The standard approach relies on word profiles derived from articles citing a cluster of co-cited articles. ... Word-profile approaches have drawbacks. First, word profiles may not converge to a focused message. Analysts and users will make a substantial amount of sense-making efforts to synthesize a diverse range of word profiles. Second, cluster labels based on aggregating word profiles tend to be too broad to be useful. In practice, many users would be interested in not only the most commonly used terms but also terms that can lead to profound changes. Terms associated with an emerging trend could be overshadowed by a broader and more persistent theme.
In CiteSpace II, a current research front is identified based on such burst terms extracted from titles, abstracts, descriptors, and identifiers of bibliographic records. These terms are subsequently used as labels of clusters in heterogeneous networks of terms and articles.
CiteSpace II makes it easier for users to identify pivotal points. In addition to inspecting salient visual attributes, the user easily can see nodes with high betweenness centrality (Freeman, 1979).
The procedure of using CiteSpace II is described in the following steps, 
(1) Identify a knowledge domain using the broadest possible term.
(2) Data collection
(3) Extract research front terms: CiteSpace II first collects n-grams, or terms, from titles, abstracts, descriptors, and identifiers of citing articles in a dataset. The present study used single words or phrases of up to four words. ... Research-front terms are determined by the sharp growth rate of their frequencies.
(4) Time slicing
(5) Threshold selection
(6) Pruning and merging: Pathfinder network scaling is the default option in CiteSpace II for network pruning (Chen, 2004; Schvaneveldt, 1990).
(7) Layout
(8) Visual inspection
(9) Verify pivotal points
We demonstrate the new features of CiteSpace with case studies of two research fields: mass-extinction research (1981–2003) and terrorism research (1990–2003).
Mass-extinction research (1981–2003).
The input data for CiteSpace II were retrieved from citation index databases via the Web of Science based on a topic search for articles published between 1981 and 2003 on mass extinction. The scope of the search included four topic fields in each bibliographic record: title, abstract, descriptors, and identifiers. The search was limited to articles in English only.
The resultant dataset contains a total of 771 records.
A total of 333 research-front terms were detected from the four topic fields of these records.
Terrorism research (1990–2003).
The terrorism research (1990–2003) dataset consists of 1,776 records resulted from a topic search on terrorism in the Web of Science.
A total of 1,108 research-front terms were found.
The fully integrated representation of research fronts and intellectual bases in the same network visualization has three practical advantages.
First, using surged topical terms rather than the most frequently occurring title words is particularly suitable for detecting emerging trends and abrupt changes. In visualized networks, research-front terms are explicitly linked to intellectual-base articles. This design presents a compact representation of the duality between a research front and its intellectual base.
Second, research-front terms naturally lend themselves to be used as labels of specialties.
Third, it overcomes a common drawback of word-profile-based labeling approaches. Aggregated word profiles may not converge to an intrinsic focus. Terms selected based on sudden increased popularity measures are particularly suitable to characterize a current research front.
The Pathfinder algorithm extracts the most salient patterns from a network, but it does not scale well. CiteSpace II implements a concurrent version of the algorithm. The concurrent Pathfinder algorithm has substantially optimized the network scaling module, although it still took 6,000 seconds to process 14 networks and merge them into a 1,704-node network.
In conclusion, the new features introduced to CiteSpaceII for detecting and visualizing emerging trends and abrupt changes in a field of research have produced promising and encouraging results. The major findings are that
• the surge of interest is an informative indicator for a new research front;
• using heterogeneous networks of terms and articles provides a comprehensive representation of the dynamics of a specialty;
• research-front terms are informative cluster labels;
• citation tree-ring visualizations are visually appealing and semantically interpretable;
• betweenness centrality metrics identify semantically valid pivotal points.

2013年12月19日 星期四

McCain, K. W. (1990). Mapping authors in intellectual space: a technical overview. Journal of the American Society for Information Science, 41(6), 433-443.

McCain, K. W. (1990). Mapping authors in intellectual space: a technical overview. Journal of the American Society for Information Science, 41(6), 433-443.

vis_paper

本論文說明作者共被引分析(author cocitation analysis, ACA)的進行步驟與相關技術,ACA的分析流程包括1)選取即將分析的作者集合、2)取得作者的共被引次數、3)建立原始的作者共被引矩陣、4)利用原始共被引矩陣計算相關係數,每一對作者之間以他們與其他作者共被引次數分布的相似程度作為他們之間的接近值、5)對接近值矩陣進行叢集分析(cluster analysis)、多維尺度分析(multi-dimensional scaling, MDS)和因素分析(factor analysis),產生視覺化圖形、6)解釋與驗證。
Within a given map, the proximity of points representing authors reflects their perceived similarity on some dimension. By examining the distribution of authors and author clusters within the two- or three-dimensional “intellectual space” of a mapped display, other aspects of structure can be described. Clusters of points can be identified with subject areas, research specialties, schools of thought, shared intellectual styles, or temporal or geographic ties. In a factor analysis, factor loadings may demonstrate the breadth or concentration of various authors’ scholarly contributions.
A common sequence of steps in author cocitation analysis is as follows,
1) Selection of the author set: One relatively objective way to identify potentially well-cited authors is to choose those who have many page references in a text, monograph, or collection of review articles.
2) Retrieval of cocited author counts
3) Compilation of  raw cocitation matrix
4) Conversion of the raw data matrix to a matrix of proximity values: The creation of a correlation matrix has at least two major advantages. First, for any given pair of authors, the correlation coefficient functions as a measure, not just of how often that pair of authors were cocited (the raw frequency count), but of how similar their “cocitation profiles” are. ... The correlation coefficient also removes differences in “scale” between authors who are highly cited and those who have similar profiles but are less frequently cited overall (Kerlinger, 1973). ... The correlations are defined as measures of similarity: the higher the positive correlation, the more similar two authors are in the perceptions of citers.
5) Approaches to multivariate analysis have been used to display the inter-author relationships in the similarities matrix:
a) In ACA, cluster analysis is used to group authors so as to provide insights into the intellectual organization of a given field. ... The two most popular approaches to cluster formation are called “hierarchical agglomerative” vs. “iterative partitioning” ... ACA research has tended to use the agglomerative clustering approach. The hierarchical agglomerative methods can use the correlation matrix as similarity measures among the authors. Authors are paired, an author is joined to an existing cluster, or two clusters are fused based on their similarity.
b) Multidimensional scaling (MDS) requires as input the same matrix of similarities or dissimilarities among objects as cluster analysis, and the two are often used together. MDS is a set of techniques used to create visual displays- maps -from proximity matrices, so that the underlying structure within a set of objects can be studied. In ACA, the major uses of multidimensional scaling are two-fold -to provide an information-rich display of the cocitation linkages and to identify the salient dimensions underlying their placement. ... Authors heavily cocited (because of their common subject or methodological interests) appear grouped in space. Authors with many links to others tend to be in central positions, while authors weakly linked, or with a few focused ties, will be placed in the periphery. In this way, “central” and “peripheral” research specializations, schools of thought, or other intellectual groupings can easily be seen, ... Dimensions are interpreted based on examination of the author and cluster placements.  ... The stress value reported for each solution (usually Kruskal’s Stress I or Stress II) and the proportion of variance explained (R Square in ALSCAL) are indicators of the overall “goodness of fit” of that point configuration.
c) Factor analytic techniques may be used to complement MDS and clustering displays. ... Essentially, they attempt to “explain” the interrelationships observed among the original variables through the creation of a much smaller number of “derived” variables or factors. In ACA, a factor is interpreted by the subset of authors loading on it - i.e., making substantial contributions to its construction. Essentially it reveals their underlying subject matter, as perceived by citers. ... ACA most commonly uses a principal components analysis, with an orthogonal (varimax) rotation of the extracted factors.
6) Interpretation and Validation: In ACA, interpretation and validation of results generally interact. Interpretation relies on discovering what the author clusters, factors, and map dimensions represent in terms of scholarly contributions, institutional or geographic ties, intellectual associations, and the like

2013年4月20日 星期六

Boyack, K. W., Klavans, R., & Börner, K. (2005). Mapping the backbone of science. Scientometrics, 64(3), 351-374.

Boyack, K. W., Klavans, R., & Börner, K. (2005). Mapping the backbone of science. Scientometrics, 64(3), 351-374.

information visualization

一般的科學映射流程 [7] 包含以下的五個步驟:1) 選擇適合的資料來源;2) 選擇分析的單位,並從選擇的來源中抽取需要的資料;3) 選擇適合的相似程度(similarity)測量方式,並計算相似值;4) 利用定位(ordination)與叢集(clustering)演算法產生資料映射圖;5)根據映射圖結果進行探索,以解答研究的問題。過去的期刊映射研究大多使用期刊共被引資料的Pearson相關係數做為期刊間相似程度的計算方式,並利用多維尺度法(multidimensional scaling, MDS)進行定位來產生圖形[11~16]。Leydesdorff [17,18]也使用多維尺度法作為映射方法,但是他使用期刊之間的引用資料測量期刊間的相似程度。Leydesdorff 也曾經進一步應用期刊之間引用資料的Pearson相關係數測量期刊間的相似程度,並且使用Pajek程式 [27]分別繪製SCI和SSCI期刊的映射圖[22, 23]。此外,Campanario [19]利用自組織映射圖(self-organizing map)作為定位的方式。Tijssen and van Leeuwen [20]則利用期刊內容映射(journal content mapping),使得他們的研究可以包含非ISI資料庫內的期刊。

本研究對7121種ISI的SCI和SSCI資料庫裡的期刊,產生代表所有科學結構的映射圖,分析產生圖形的結構準確性(structural accuracy),同時也探討映射結果的局部準確性(local accuracy),後者所指的是屬於相同次學科(subdiscipline)的期刊能夠被群聚在一起(Klavans and Boyark, 2006),而前者所指的是彼此相互引用的期刊群聚在映射到圖形上時也有相互鄰近,也就是本研究所認為的科學的主幹(backbone of science)。

在這項研究中,利用VxOrd [32]演算法進行資料定位,以k-means法進行叢集,以八種方式測量期刊之間的相似程度,其中五種是利用期刊之間的引用(inter-citation)資料為基礎,包括原始次數、Cosine指標、Jaccard指標、Pearson相關係數和由Pudovkin and Garfield [25]提出的平均關連因數(average relatedness factor);另外三種是以共被引(cocitation)資料為基礎,包括原始資料、Pearson相關係數和作者等人[1]提出的K50指標。

本研究利用期刊在ISI的主題分類(subject categories)評估與比較各種相似程度測量方式所產生的映射圖結果。首先在局部準確性的評估上,本研究認為某一對期刊間如果具有較高的相似程度便應在同一主題分類中,映射圖上同一叢集的期刊也應屬於同一主題分類。[1]的研究結果發現經過VxOrd 演算法進行資料定位後可以增加局部準確性,並且以期刊之間的引用資料為基礎並經過正規化的四種指標表現相當,而以Cosine指標為最佳。本研究則以Gibbons and Roth [37] 所提出的交互資訊(mutual information)的檢測法來比較各種映射圖結果的結構準確度,同樣以ISI的205個主題分類為參考的分類標的。檢測的結果除了共被引的原始資料的分群結果不理想之外,其餘的各種相似程度測量方式所得的分群結果大致相當,並且以期刊之間引用資料為基礎的Pearson相關係數得到最好的結果,但是在200到250個分群時,以期刊之間引用資料為基礎的Jaccard指標也有相當的分類結果。

最後以局部準確性、結構準確性、規模可擴展性(scalability)和叢集結果的可判讀性(readability)對以期刊之間引用資料為基礎的五種方式和以期刊之間共被引資料為基礎的三種方式分別進行比較。以期刊之間共被引資料為基礎的三種方式而言,利用K50指標為相似程度測量方式所得到的結果在結構準確性上與Pearson相關係數所得到的結果大約相當,但前者具有較好的規模可擴展性和局部準確性,並且其呈現的結果在叢集大小與分布位置上都較為均衡。另一方面,在以期刊之間引用資料為基礎的五種方式,Cosine指標、Jaccard指標、Pearson相關係數等三種在規模可擴展性和可判讀性較其他兩種為佳,並且Jaccard指標所得到的映射圖有較高的結構準確性。

This paper presents a new map representing the structure of all of science, based on journal articles, including both the natural and social sciences. ... Eight alternative measures of journal similarity were applied to a data set of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes. For each journal similarity measure we generated two-dimensional spatial layouts using the force-directed graph layout tool, VxOrd.

By accuracy, we mean that journals within the same subdiscipline should be grouped together, and groups of journals that cite each other should be proximate to each other on the map. The first results from this effort, dealing with local accuracy, appeared recently. By contrast, this paper focuses on structural accuracy and characterization of the map defining the structure or backbone of science.

Published journal-based maps have typically been focused on single disciplines, and have used a Pearson correlation on co-citation counts with multidimensional scaling (MDS). [11~16] Other discipline-level studies not using the Pearson/MDS technique include the use of relative inter-citation counts with MDS by Leydesdorff [17,18], the use of a self-organizing map by Campanario [19], and the work by Tijssen and van Leeuwen to include non-ISI journals in their maps using journal content mapping. [20]

Leydesdorff has used the 2001 JCR data to map 5,748 journals from the Science Citation Index (SCI) [22] and 1,682 journals from the Social Science Citation Index (SSCI) [23] in two separate studies. In both studies Leydesdorff uses a Pearson correlation on citing counts as the edge weights and the Pajek program for graph layout, progressively lowering thresholds to find articulation points (i.e., single points of connection) between different network components. These network components are his journal clusters. The only potential drawback to this solution is that as thresholds are lowered, newly identified small components (presumably two or three journals each) are dropped from the solution space, so that the total number of journals comprising Leydesdorff's clusters is substantially less than the number in the original set.

An alternative to using journals to map the structure of science has recently been investigated by Moya-Anegón and associates [9] to good effect. Using 26,062 documents with a Spanish address from the year 2000 as a base set, they used co-cited ISI category assignments to create category maps. Their highest level map shows the relative positions, sizes and relationships between 25 broad categories of science in Spain.

The general process followed by most practitioners for creating knowledge domain maps has been explained in detail elsewhere. [7] This process can vary slightly depending upon the specific research question, but typically contains the following steps: 1) selection of an appropriate data source, 2) selection of a unit of analysis (e.g. paper, journal, etc.) and extraction of the necessary data from the selected source, 3) choice of an appropriate similarity measure and calculation of similarity values, 4) creation of a data layout using a clustering or ordination algorithm, and 5) exploration of the map based on the data layout as a means of answering the original research questions. Here, we add another step after 4) - statistical validation - that allows us to choose the similarity measure that produces the most accurate map.

Based on these considerations, we obtained the complete set of 1.058 million records from 7,349 separate journals from the combined SCI and SSCI files for the year 2000. Of the 7,349 journals, analysis was limited to the 7,121 journals that appeared as both citing and cited journals. ... Journal inter-citation frequencies were directly counted from the citing and cited journal information in these 16.24 million reference pairs. The resulting journal-journal inter-citation frequency matrix was extremely sparse (98.6% of the matrix has zeros). ... While there was a great deal more cocitation frequency information, the journal-journal co-citation frequency matrix was also sparse (93.6% of the matrix has zeros).

For the purpose of map validation we also retrieved the ISI journal category assignments. For the combined SCI and SSCI, there were a total of 205 unique categories. Including multiple assignments, the 7,121 journals were assigned to a total of 11,308 categories, or an average of 1.59 categories per journal.

The five inter-citation measures include one unnormalized measure, raw frequency (IC-Raw); and four normalized measures, Cosine (IC-Cosine), Jaccard (IC-Jaccard), Pearson’s r (IC-Pearson), and the recently introduced average relatedness factor of Pudovkin and Garfield [25] (IC-RFavg).

The three co-citation measures include one unnormalized measure, raw frequency (CC-Raw); the vector-based Pearson’s r (CC-Pearson), and a new normalized frequency measure [1] that we call K50 (CC-K50). This new measure, K50, is simply a cosine-type value minus an expected cosine value. Ei,j is the expected value of Fi,j, and varies with the row sum, Sj, thus K50 is asymmetric and Eij<>Eji . Subtraction of an expected value component tends to accentuate ‘higher than expected’ relationships between two small journals or between a small and a large journal, and discounts ‘lower than expected’ relationships between large journals. We thus expect the K50 measure to do a better job than other measures of accurately placing small journals, and to reduce the influence of large and multidisciplinary journals on the overall map structure.

The most commonly used reduction algorithm is multidimensional scaling; however, its use has typically been limited to data sets on the order of tens or hundreds of items.

Factor analysis is another method for generating measures of relatedness. In a mapping context, it is most often used to show factor memberships on maps created using either MDS or pathfinder network scaling, rather than as the sole basis for a map. Yet, factor values can be used directly for plotting positions. For instance, Leydesdorff [23] directly plotted factor values (based on citation counts) to distinguish between pairs of his 18 factors describing the SSCI journal set.

Layout routines capable of handling these large data sets include Pajek, [27] which has recently been used on data sets with several thousand journals by Leydesdorff, [22,23] and which is advertised to scale to millions of nodes; self-organizing maps, [28] which can scale, with various processing tricks, to millions of nodes, [29] and the bioinformatics algorithm LGL, [30] capable of dealing with hundreds of thousands of nodes, which uses an iterative layout as well as data types and algorithms from the Boost Graph Library. [31]

We chose to use VxOrd, [32] a force-directed graph layout algorithm, over the other algorithms mentioned, for several reasons. VxOrd improves on a traditional force-directed approach by employing barrier jumping to avoid trapping of clusters in local minima, and a density grid to model repulsive forces. Because of the repulsive grid, computation times are order O(N) rather than O(N(square)), allowing VxOrd to be used on graphs with millions of nodes. VxOrd also applies edge cutting criteria, which leads to graph layouts exhibiting both local (orientation within groups) and global (group-to-group) structure. The combination of the initial node and edge structure and cutting criteria thus determine the number, size, shape, and position of natural groupings of nodes.

Validation of science maps is a difficult task. In the past, the primary method for validating such maps has been to compare them with the qualitative judgments made by experts, and has been done only for single-discipline-scale maps (see the background section of Klavans & Boyack [1] for more discussion).

A more pragmatic approach is to use the ISI journal classifications to evaluate the validity of the journal similarity measures and the corresponding maps. The ISI journal classification system, while it does have its critics, is based on expert judgment and is widely used. In principle, users would expect that pairs of journals with high similarity should be in the same ISI category. Journals in the same cluster of a journal mapping should have the same ISI category assignments. These assumptions are used to validate and compare the eight different similarity measures and corresponding graph layouts or maps.

In our previous work with the current data set, and the same eight similarity measures and maps from Figure 1, we investigated local accuracy and the effects on accuracy of reducing dimensionality with VxOrd [1] using the ISI category assignments as a reference basis. We found that, counterintuitively, use of VxOrd algorithm to convert similarities to map positions actually increased local accuracy. We also found that four of the inter-citation measures had roughly comparable local accuracy at 95% journal coverage, and recommended the IC-Cosine measure as the best overall measure.

In this work we focus on structural accuracy or the validity of the global structure of the solution space. To make quantitative comparisons of our eight maps of science, we implement a mutual information method recently used to distinguish between gene clustering algorithms. [37] This mutual information method requires a reference basis, for which we use the ISI journal category assignments.

To employ the method of Gibbons and Roth [37] we need to do a clustering of each of the maps. VxOrd gives (x,y) coordinate positions for each node, but does not assign cluster numbers to the nodes. Thus, k-means clustering was applied to each of the maps in Figure 1. Other clustering methods (e.g. linkage or density-based clustering) could have been used.

The CC-Raw map clearly performs the worst. The Z-scores for all other measures are near or above a value of 350, indicating that all of these measures give maps that are far from random. The IC-Pearson map gives the highest Z-score over nearly the entire range of cluster solutions. It is only at the higher end, from 200 through 250 clusters, that the IC-Jaccard map has a Z-score comparable to that of the IC-Pearson.

Hence, based on Z-scores it is likely that any of the six would be a suitable choice as the basis for an accurate map of science.

For a co-citation-based map, the CC-K50 measure is a clear winner for several reasons. Although the Z-score for the CC-K50 is nearly identical to that of the CC-Pearson, the K50 measure is scalable to much larger numbers of nodes, while the Pearson is a full N(square) calculation, and cannot easily scale much higher than the 7000 nodes used here. The CC-K50 map is a visually well-balanced map with a good distribution of cluster sizes and positions (see Figures 1 and 3). By contrast, the CC-Pearson map appears very stringy; clusters are very dense with less visual differentiation between disciplines, and thus not as suitable for presentation. The CC-K50 map also has a higher degree of local accuracy. [1]

Of these three, IC-Cosine, IC-Jaccard, and IC-Pearson, we choose to further characterize the IC-Jaccard as our best map due to its slightly higher Z-score, realizing that the Cosine map is in a virtual dead heat statistically, and the Pearson map only somewhat less in local accuracy.

Differences such as these between the maps at the discipline level are likely due to fine-scaled differences between the co-citation and inter-citation patterns. Yet, the overall consistency between the co-citation and inter-citation-based maps of science suggests the general structure described here is robust.

Figure 6 shows the clear distinction between two main areas within the LIS discipline. Although there are relationships between journals in the two clusters, the dominant relationships (darkest edges) are within clusters. The journals in the cluster at the upper left all focus on libraries and librarians and their work, while those in the cluster at the lower right are all focused on advances in information science.

Eight different similarity measures were calculated from the combined SCI/SSCI data and the resulting journal-journal similarity matrices were mapped using VxOrd. The eight maps were then compared based on two different accuracy measures, the scalability of the similarity algorithm, and the readability of layouts (clustering).

2013年4月8日 星期一

Yan, E., Ding, Y., Milojević, S., & Sugimoto, C. R. (2012). Topics in dynamic research communities: An exploratory study for the field of information retrieval. Journal of Informetrics, 6(1), 140-153.

Yan, E., Ding, Y., Milojević, S., & Sugimoto, C. R. (2012). Topics in dynamic research communities: An exploratory study for the field of information retrieval.Journal of Informetrics6(1), 140-153.

information visualization

本研究整合社群偵測(community detection) (Clauset, Newman, & Moore, 2004)和主題確認(topic identification) (Tang et al., 2008)等兩個演算法,藉以探討研究社群與研究主題之間的彼此交織(interwoven)且共同演化(co-evolving)的互動關係。目前已經有相當多研究分別對於科學研究的研究社群和主題確認進行探討,前者例如利用作者的合著網絡(coauthorship networks)進行社群偵測來找出社群間互動模式(patterns of community interactions)的研究(Girvan & Newman, 2002; Richardson, Mucha, & Porter,2009; Pepe & Rodriguez, 2010),後者的研究則有利用詞語的機率分布做為研究領域上各個主題的統計模型(model)(Blei, Ng, & Jordan, 2003)。在整合兩種研究方向上,Racherla and Hu (2010)對於作者合作的主題進行單一或多樣性的探討,Steyvers, Smyth, Rosen-Zvi, and Griffiths (2004)將作者的統計模型描述為主題上的機率分布,稱為作者-主題模型(Actor-Topic Model)。根據作者-主題模型,McCallum, Corrada-Emmanuel, and Wang (2004)進一步提出作者-接收者-主題模型(Actor-Recipient-Topic Model),將主題描述為詞語的multi-normal分布,而每一對作者和接收者則是在主題上的分布。Tang et al. (2008)則是將作者-主題模型加以擴大,提出作者-研討會-主題模型(Actor-Conference-Topic Model),以機率模型同時描述文章內容、作者興趣以及研討會(期刊)。
本研究以上述的研究為基礎,以2001年到2007年的資訊檢索(information retrieval, IR)領域為研究案例,將這期間出版的論文分為三個時期(2001-2003, 2004-2005, 2006-2007),找出每個時期在IR領域的主要研究社群和主題,並探討主題之間以及每一個社群與主題之間的相關性。本研究使用Clauset et al.’s (2004)提出的社群偵測演算法從三個時期的作者合著網絡的最大相連成分上發現主要的研究社群,Clauset et al.’s (2004)演算法屬於目前常用的模組性(modularity)最大化的社群偵測演算法,每一個時期找出最大的10個社群。確認研究主題的方式則是使用Tang et al., (2008)的ACT模式,使得每一個作者在每個主題上有一個機率分布的模型。每個時期確認出10個最主要的主題,將每個時期的主題與同一時期的其他主題利用Cosine對它們的統計模型進行相似性的比較,結果發現除了這些主題彼此間的相似性都很低,顯示ACT模式能夠很好地區辨出不同的主題。比較不同時期的主題可以發現主題間的延續關係,主題和它們的延續者之間的相關性較高。研究結果指出較普遍的主題具有較多的延續主題,較不普遍的主題則僅有一個延續主題,甚至沒有。在比較研究社群和主題時,發現規模較大的社群,因為包含較多的作者而涉及較多元的研究興趣,所以很難有集中的研究主題;反之,規模較小的社群其研究主題較集中。從這個研究裡可以看出作者傾向於和有相似研究專長且發表相近主題論文的其他作者進行合作。最後,這個研究也指出某些社群致力於相當獨特、其他社群相當少涉及的研究主題,例如生物醫學(biomedical)相關的主題。

This paper examine show research topics are mixed and matched in evolving research communities by using a hybrid approach which integrates both topic identification and community detection techniques. Using a data set on information retrieval(IR) publications,two layers of enriched information are constructed and contrasted: one is the communities detected through the topology of coauthorship network and the other is the topics of the communities detected through the topic model.

The complexity of scholarly data has led to a growing interest in applying probabilistic models to identify topics from documents. A topic represents an underlying semantic theme and can be informally approximated as an organization of words and can be formally operationalized as a probability distribution over terms in a vocabulary (Blei & Lafferty, 2007)

Topic models are the latest advancement in this vein of research (e.g. Blei, Ng, & Jordan, 2003). Topic models provide useful descriptive statistics for a collection of scholarly data, thus making it easier for scholars to navigate academic documents. The outcomes of topic models are probability distributions of words or publications for each topic (e.g. Blei et al., 2003); however, they provide no information on which community contributes to a certain topic or how topics are developed by communities.

Research communities can be detected using community detection methods to group actors, such as authors and journals, with the goal of identifying patterns of community interactions.

Leskovec, Lang, Dasgupta, and Mahoney (2008) used the concept of “conductance” to capture a community: a good community should have small conductance, i.e. “it should have many internal edges and few edges pointing to the rest of the network” (p. 4).

A decisive advance in community detection was made by Newman and Girvan (2004),who introduced a quantitative measure for the quality of partitioned communities, a.k.a. the modularity.

In reality, communities and topics are not disconnected; on the contrary, communities and topics are interwoven and co-evolving: that is, a research community can carry several topics, and a topic can consist of different collaboration groups (Li et al., 2010a). Therefore, in order to study the interdisciplinary nature of science, it is necessary to integrate the two threads of research on community detection and topic identification, and utilize them to understand the dynamic interactions between topics and communities.

Racherla and Hu (2010) constructed a topic similarity matrix by assigning a predefined research topic to each document and its authors, and using authors’ collaboration information to link topics. They found that authors not only collaborate on the same research topics but also collaborate on varied research topics.

Upham, Rosenkopf, and Ungar (2010) developed an iterative clustering scheme that produces high-quality dynamic clusters over time. Using such an approach, twenty-one research communities were detected in the information science and technology area. Innovation performance was then quantified by various parameters and measured for each of these clusters.

Pepe and Rodriguez (2010) conducted an in-depth study of a small collaboration network of researchers in the area of sensor networks and wireless technologies. They adopted the notion of discrete assortativity coefficient to evaluate the collaboration pattern in this network. They found that its collaboration has become more intra-institutional and more inter-disciplinary.

Built upon previous endeavors on graph partitioning, Girvan and Newman (2002) proposed an algorithm that uses edge betweenness to identify the boundaries of communities. They applied the method to a scientific collaboration network at the Santa Fe Institute, and identified several densely connected communities. They found that scientists are grouped together either by a similar research topic or by a similar research methodology, where the latter situation may be an indication of interdisciplinary work.

The Girvan–Newman algorithm is computationally time demanding and is optimized into a more efficient algorithm (Clauset, Newman, & Moore, 2004). The new algorithm incorporated modularity, now becoming a standard measure to evaluate community structures. For instance, Richardson, Mucha, and Porter (2009) found their spectral graph-partitioning algorithm can yield higher-modularity partitions. They applied their method to a coauthorship network of network scientists and found three well-known research centers in network science. However, from their findings, it is unclear whether the three locations also form three distinct research topics or how the research centers are connected via topics.

Similar to the methods used in detecting author communities, scholars working on identifying topics have used methods such as multidimensional scaling (e.g. White & McCain, 1998), k-means (e.g. Yan, Ding, & Jacob, in press), modularity-based clustering techniques (e.g. Van Eck & Waltman, 2010), and hybrid approaches (e.g. Janssens, Glänzel, & De Moor, 2008).

Upham and Small (2010), for instance, gave a good quantitative definition of growing, shrinking, stable, emerging, and exiting research fronts.

Traditionally, the research instruments they utilize are mainly co-occurrence networks, for instance, author co-citations networks (White & McCain, 1998), document co-citation networks (Klavans & Boyack, 2011; Small, 1973; Upham & Small, 2010), journal co-citation networks (Ding, Chowdhury, & Foo, 2000a), or co-word relations (Ding, Chowdhury, & Foo, 2000b; Milojevic, Sugimoto, Yan, & Ding, 2011).

Boyack and Klavans (2010) examined several types of scholarly networks, including a cocitation network, a bibliographic coupling network, and a citation network, in the interest of selecting the network that can represent the research front in biomedicine. They used within-cluster textual coherence and grant-to-article linkage indexed by MEDLINE as accuracy measurements and found that the bibliographic coupling-based citation-text hybrid approach, an approach that couples both references and words from title/abstract, outperformed other approaches.

Janssens, Glänzel, and De Moor (2007, 2008) proposed a novel hybrid approach that integrates two types of information, citation (in the form of a term-by-document matrix) and text (in the form of a cited references-by-document matrix). Noticing that the weighted linear combinations may “neglect different distributional characteristics of various data sources” (p. 612), the authors developed a new approach named Fisher’s inverse chi-square method. This method can effectively combine matrices with different distributional characteristics. They found the hybrid approach outperformed the text-only approaches by successfully assigning papers into correct clusters.

Followed by the tradition of data mining and knowledge discovery, topic models have gained great popularity among computer scientists in recent years. One wellknown topic model is the Probabilistic Latent Semantic Indexing (pLSI) model proposed by Hofmann (1999). Built on pLSI, Blei et al. (2003) introduced a three-level Bayesian network, called Latent Dirichlet Allocation (LDA). In topic models, topics are modeled as a probability distribution over terms in a vocabulary.

Steyvers, Smyth, Rosen-Zvi, and Griffiths (2004) proposed an unsupervised learning technique for extracting both the topics and authors of documents. In their Author-Topic model, authors are modeled as probability distributions over topics.

McCallum, Corrada-Emmanuel, and Wang (2004) presented the Author-Recipient-Topic (ART) model, a directed graphical model which conditions the per-message topic distribution jointly on both the author and individual recipients. In ART model, each topic is modeled as a multi-nomial distribution over words, and each author–recipient pair is modeled as a distribution over topics.

The Author-Conference-Topic (ACT) model, proposed by Tang et al. (2008), further extended Author-Topic model to include conference/journal information. The ACT model utilizes probabilistic models to model documents’ contents, authors’ interests, and also conference/journal simultaneously.

To understand the interaction between research communities and research topics, there is a need to incorporate both community detection and topic modeling approaches. For instance, Zhou, Ji, Zha, and Lee Giles (2006) and Zhou, Manavoglu, Li, Lee Giles, and Zha (2006) proposed two generative Bayesian models for semantic community detection in social networks by combining probabilistic modeling with community detection algorithms. Their method was able to detect the communities of individuals and meanwhile provide topic descriptions to these communities.

Information retrieval (IR) was chosen as the target domain. Papers were collected from Scopus for 2001–2007 (inclusive)... Time slices were set as 2001–2003, 2004–2005, and 2006–2007 so that each slice has similar number of authors, thus
providing comparable networks. Authors in the largest component (LC) were finally selected to form the coauthorship networks (Table 1).

Clauset et al.’s (2004) method was implemented to the coauthorship networks for each time period. The modularity for weighted networks can be calculated as (Clauset et al., 2004): ... Formula (1) is the fraction of within-community edges minus the expected value of the same degrees of vertices randomly connected between the vertices

An extended stop word list is used to exclude common words in IR, including information, retrieval, system, search, and model. The ACT model (Tang et al., 2008) was used to detect topics. In the ACT model, each author is associated with a multi-nomial distribution over topics and words in a paper and the conference stamp is generated from a sampled topic.

In this way, the posterior distribution of topics depends on three modalities: authors, words, and conferences (or journals). The model begins with the joint probability of the whole data set, and then using the chain rule, the posterior probability of sampling the topic and author for each word can be obtained.

The next step was to overlay research topics for the detected communities. The procedures were: (1) search and collect publications for all authors in the top ten communities in each time slice; (2) apply the ACT model to publications of each
time slice with the number of topics set at ten; (3) generate an topic-author distribution (P(topic | author)) using the ACT model where each author obtains a topic distribution vector (for author i: ai = (t1, t2, . . . , t10)), and set up a threshold and replace those probabilities that below the average 0.1 (1/10) to 0; by doing so, the insignificant probabilities will not be counted and will not add noise to the community similarity calculation; (4) extract and average the topic distributions for authors of a community where the mean is considered as the community’s topic distribution vector, and then normalize the vector so that the sum of each vector is one; (5) calculate cosine similarities for communities.

An increasing number of words have been added to the knowledge domain of IR over time: from 3785 in 2001–2003 to 9794 in 2006–2007, indicating an expanded research scope of IR scholars. Around half of the words used in the earlier period are inherited by the next period—the other half is abandoned. In addition, 10% of the words in 2001–2003 were not mentioned in 2004–2005 but regained attention in 2006–2007.

For topics belonging to the same time period (the three blocks located on the diagonal line), most topics have low similarities with other topics. It is a good sign in that the ACT model has successfully identified distinguishable topics.

For topics belonging to different time periods, it can be found that some topics have evident successors (bright squares) while other topics fail to proceed into the next time period (dark squares). In addition, it can also be found that topics with high popularities tend to have multiple successors and topics with low popularities tend to have only one or none successor.

Topic popularity is predicated on the ACT model. The underlying assumption is that if the words belonging to a certain topic occur more frequently, then this topic has high popularity. Since ten topics are set, a topic popularity of 0.1 means this topic has an average popularity. A value above 0.1 suggests a “hot” topic and a value below 0.1 suggests a “cold” topic.

Topics in high popularities are well connected: the top five topics in each time period have predictors and/or successors. However, topics in low popularities are loosely connected, suggesting that they did not receive continuous attention.

Communities of smaller sizes tend to have evident topical concentrations, which is understandable as communities of larger sizes are more likely to involve scholars with diverse research interests. For example, in 2001–2003, Community 7 is specialized in Topic 10, Community 9 is specialized in Topic 7, and Community 10 is specialized in Topic 6; comparatively, the top three communities in 2004–2005 and 2006–2007 did not yield evident topical concentrations.

In regard to topics, most topics are associated with at least one community. For example, Topic 9 in 2004–2005 is studied by Community 4, and Topic 4 in 2006–2007 is studied by Community 10. The results indicate that authors are more inclined to collaborate with others who have similar expertise and publish papers on similar topics. In addition, smaller communities tend to have relatively distinct research topics.

A few communities (Community 7, 9, and 10 in 2001–2003; Community 4 and 10 in 2005–2006; Community 6 and 10 in 2006–2007) concentrates on relatively unique topics and has lower level of topical similarity with other communities,
especially for biomedical related topics, such as Community 4 in 2004–2005 is highly specialized in “database-protein-gene-expression-mining”, Community 10 in 2004–2005 is highly specialized in “medical-health-clinic-systematic-biomedical”, and Community 10 in 2006–2007 “application-memory-optical-remote-imaging”.

2013年4月4日 星期四

Cobo, M. J., López-Herrera, A. G., Herrera-Viedma, E., & Herrera, F. (2011). An approach for detecting, quantifying, and visualizing the evolution of a research field: A practical application to the fuzzy sets theory field. Journal of Informetrics, 5(1), 146-166.

Cobo, M. J., López-Herrera, A. G., Herrera-Viedma, E., & Herrera, F. (2011). An approach for detecting, quantifying, and visualizing the evolution of a research field: A practical application to the fuzzy sets theory field. Journal of Informetrics5(1), 146-166.

information visualization

書目計量學的兩個主要研究方向:一為成效分析(performance analysis),也就是根據書目資料評估國家、大學、研究者等群體的表現與他們的影響力(Noyons, Moed, & van Raan, 1999; van Raan, 2005a);另一為科學對映(science mapping),是指利用科學對映圖呈現科學研究的結構與動向 (Börner, Chen, & Boyack, 2003; Noyons, Moed, & Luwel, 1999)。此類研究方向通常以作者、參考文獻或是詞語等物件在論文裡的共現(co-occurrence)關係為基礎,進行分析物件的叢集,使得在同一叢集內的物件之間彼此具有較強於叢集間物件的關係,所得到的叢集結果彼此間的關係便可以表現出科學研究的結構。一般而言,以共同出現在論文的關係所獲得的作者叢集可以代表科學研究的社會結構(social structure),參考文獻代表科學研究的知識基礎(intellectual base),其叢集與其間的關係便是呈現出相關的知識結構(intellectual structure),另外詞語的叢集則可以表示科學研究的主題(theme),呈現出相關的認知結構(cognitive structure)。將分析的時間區分成多個時段,在多個時段間彼此相關連的主題,也就是這些主題之間有相當多重複的詞語,為了分析這些延續性的主題之間的演變(evolution),本研究將這些主題形成的序列稱為主題區域(thematic area)。過去對於成效分析的研究,其方法著重於對群體的表現,較少以認知的方式對研究領域內的特定主題以及主題區域進行生產力與影響力的評估。換言之,便是較少整合上述的兩個研究方向。
為了上述的目的,本研究以下列的步驟進行科學對應與成效分析:1)利用相關強度(association strength)(van Eck & Waltman, 2007)估算詞語間的相關性,並且以簡易中央演算法(simple centers algorithm)(Coulter et al., 1998)進行叢集,確認出各個不同時段的主題;2)將叢集的結果利用網絡的密度和網絡彼此相連的強度對應到策略圖(strategic diagram)(Callon et al., 1991)上,以了解這些主體的內聚程度和外延程度;3)分析主題在連續時段上形成的主題區域,確認主題區域的起源與演變;4)測量各個主題及主題區域在書目計量學上的生產力與影響力等成效。
本研究以Fuzzy Sets Theory(FST)做為科學對映與成效分析的案例,實際操作上述的分析,研究結果包括:1)結合各種書目計量工具來分析研究領域認知結構的演化可以發現FST各階段的基礎主題與相關主題區域等重要知識,例如主題區域FUZZY-CONTROL的逐漸增長趨勢,另一主題區域FUZZY-LOGIC則走向減少。2)利用視覺化工具則可以更容易地偵測這些主題或主題區域的演化、重要性和未來走向。3)運用h-index(Alonso et al., 2009; Cabrerizo et al., 2010; Hirsch, 2005)等書目計量指標則更能夠分析各主題與主題區域的質量和影響力。

In bibliometrics, there are two main procedures: performance analysis and science mapping (Noyons, Moed, & Luwel, 1999; van Raan, 2005a). Performance analysis aims at evaluating groups of scientific actors (countries, universities, departments, researchers) and the impact of their activity (Noyons, Moed, & van Raan, 1999; van Raan, 2005a) on the basis of bibliographic data. Science mapping aims at displaying the structural and dynamic aspects of scientific research (Börner, Chen, & Boyack, 2003; Noyons, Moed, & Luwel, 1999). A science map is used to represent the cognitive structure of a research field.

The majority of these methods are mainly focused on measuring the performance of the scientific actors and little research has been carried out in order to measure the performance of given research fields in a conceptual way (specific themes or whole thematic areas). A performance analysis of specific themes or whole thematic areas can measure (quantitatively and qualitatively) the relative contribution of these themes and thematic areas to the whole research field, detecting the most prominent, productive, and highest-impact subfields.

In the case of co-citation analysis, the clusters represent groups of references that can be understood as the intellectual base of the different subfields.

On the other hand, in the case of co-word analysis, the clusters represent groups of textual information that can be understood as semantic or conceptual groups of different topics treated by the research field.

So, the detected clusters can be used with several purposes such as:
• To analyze their evolution through measuring continuance across consecutive subperiods.
• To quantify the research field by means of a performance analysis.

Cluster string (Small, 2006; Small & Upham, 2009; Upham & Small, 2010), rolling clustering (Kandylas et al., 2010) and alluvial diagrams (Rosvall & Bergstrom, 2010) have been used to show the evolution of detected clusters in successive time periods. Other authors proposed to layout the graph of a given time period taking into account previous and subsequent ones (Leydesdorff & Schank, 2008), or to pack synthesized temporal changes into a single graph (Chen, 2004; Chen et al., 2010).

Strategic diagrams (Callon et al., 1991), self-organizing maps (Polanco, Franc¸ ois, & Lamirel, 2001), heliocentric maps (Moya-Anegón et al., 2005), geometrical models (Skupin, 2009) and thematic networks (Bailón-Moreno, Jurado-Alameda, Ruiz-Banos, & Courtial, 2005; López-Herrera et al., 2009) have been proposed to show and layout the research field and its detected subfields.

To sum up, the stages carried out by our approach are:
1. To detect the themes treated by the research field by means of co-word analysis for each studied subperiod.
2. To layout in a low dimensional space the results of the first step (themes).
3. To analyze the evolution of the detected themes through the different subperiods studied, in order to detect the main general thematic areas of the research field, their origins and their inter-relationships.
4. To carry out a performance analysis of the different periods, themes and thematic areas, by means of quantitative and impact measures.

In our proposal, the process is divided into five steps: (1) collection of raw data, (2) selection of the type of item to analyze, (3) extraction of relevant information from the raw data, (4) calculation of similarities between items based on
the extracted information and (5) use of a clustering algorithm to detect the themes.

Similarities between items are calculated based on frequencies of keywords’ co-occurrences. Different similarity measures have been used in the literature, the most popular being Salton’s Cosine and the Jaccard index. In van Eck and Waltman (2009) an analysis of well-known direct similarity measures was made, concluding that the most appropriate measure for normalizing co-occurrence frequencies is the equivalence index (Callon et al., 1991; Michelet, 1988). This measure is also known as association strength (Coulter et al., 1998; van Eck & Waltman, 2007), proximity index (Peters & van Raan, 1993; Rip & Courtial, 1984), or probabilistic affinity index (Zitt, Bassecoulard, & Okubo, 2000).

Different clustering algorithms can be used to create a partition of the keywords network or graph. Recently, some authors have prosed different clustering algorithms to carry out this task: Streemer (Kandylas et al., 2010), spectral clustering (Chen et al., 2010), modularity maximization (Chen & Redner, 2010) and a bootstrap resampling with a significance clustering (Rosvall & Bergstrom, 2010).

As is described in Coulter et al. (1998), the simple centers algorithm uses two passes through the data to produce the desired networks. The first pass (Pass-1) constructs the networks depicting the strongest associations, and links added in this pass are called internal links. The second pass (Pass-2) adds to these networks links of weaker strengths that form associations between networks. The links added during the second pass are called external links.

Callon’s centrality, to be referred to as centrality henceforth, measures the degree of interaction of a network with other networks (Callon et al., 1991) ... Centrality measures the strength of external ties to other themes. We can understand this value as a measure of the importance of a theme in the development of the entire research field analyzed.

Callon’s density, to be referred to as density henceforth, measures the internal strength of the network (Callon et al., 1991) ... Density measures the strength of internal ties among all keywords describing the research theme. This value can be understood as a measure of the theme’s development.

We can find four kinds of themes (Cahlik, 2000; Callon et al., 1991; Courtial & Michelet, 1994; Coulter et al., 1998; He, 1999) according to the quadrant in which they are placed:

• Themes in the upper-right quadrant are both well developed and important for the structuring of a research field. They are known as the motor-themes of the specialty, given that they present strong centrality and high density. The placement of themes in this quadrant implies that they are related externally to concepts applicable to other themes that are conceptually closely related.

• Themes in the upper-left quadrant have well developed internal ties but unimportant external ties and so are of only marginal importance for the field. These themes are very specialized and peripheral in character.

• Themes in the lower-left quadrant are both weakly developed and marginal. The themes of this quadrant have low density and low centrality, mainly representing either emerging or disappearing themes.

• Themes in the lower-right quadrant are important for a research field but are not developed. So, this quadrant groups transversal and general, basic themes.

In a theme, the keywords and their interconnections draw a network graph, called a thematic network. Each thematic network is labelled using the name of the most significant keyword in the associated theme (usually identified by the most central keyword of the theme).

Given a thematic network, a document is called a “core document” if it has at least two keywords presented in the thematic network. If a document has only one keyword associated with the thematic network, it is called a “secondary document”. Both core and secondary documents can belong to more than one thematic network.

So, a thematic area is defined as a group of evolved themes across different subperiods. Note that, depending on the interconnections among them, one theme could belong to a different thematic area, or could not come from any.

As the themes have an associated set of documents (core documents, or secondary documents, or core documents + secondary documents), the thematic areas could also have an associated collection of documents. In this case, the documents associated with each thematic area will be ascertained through the union of the documents associated with the set of themes belonging to each thematic area.

By means of quantitative measures the productivity of the detected themes and thematic areas is analyzed, whereas qualitative measures show the (supposed) quality based on the bibliometric impact of those themes and thematic areas.
• Quantitative measures: number of documents, authors, journals and countries.
• Qualitative or impact measures: number of received citations of the documents and bibliometric indices such as the h-index (Alonso et al., 2009; Cabrerizo et al., 2010; Hirsch, 2005).

This approach combines different bibliometric tools to analyze the evolution of the cognitive structure of a research field, allowing us to discover important knowledge related to its themes and thematic areas.

In such a way, as was pointed out in Section 4.1, we discover that our approach adequately identifies the FST basic themes in each subperiod, because they achieve the highest citation scores and impacts. Additionally, as was shown in Section 4.2, we are able to identify thematic areas (see Table 7) and show their evolutionary behaviour, as with FUZZY-CONTROL whose evolution is increasing or FUZZY-LOGIC whose evolution is decreasing.

This approach is supported by different visualization tools that allow us to easily detect the themes and thematic areas and to understand their evolution, importance and likely future tendencies.

For example, we show the evolution of the FST research field in Fig. 11 and we identify that FUZZY-CONTROL is the most important thematic area with the highest impact, as is shown in Table 7.We have also concluded that FUZZY ROUGH-SETS seems to be the origin of a new thematic area.

This approach is completed by incorporating amore elaborated bibliometric index, i.e., the h-index, which allows us to better analyze the quality or impact of the themes and thematic areas. In our FST analysis, as is shown in Sections 4.1 and 4.2 we use the h-index to evaluate the impact of themes and thematic areas.

2013年3月27日 星期三

Cobo, M. J., López‐Herrera, A. G., Herrera‐Viedma, E., & Herrera, F. (2012). SciMAT: A new science mapping analysis software tool. Journal of the American Society for Information Science and Technology.

Cobo, M. J., López‐Herrera, A. G., Herrera‐Viedma, E., & Herrera, F. (2012). SciMAT: A new science mapping analysis software tool. Journal of the American Society for Information Science and Technology.

information visualization


本論文介紹科學映射分析(science mapping analysis)工具SciMAT的功能與應用。根據Börner et al., (2003)和Cobo et al., (2011b)等研究,科學映射分析的流程可以分成以下的步驟:1) 資料檢索 (data retrieval)、2) 資料前處理 (data preprocessing)、3) 網路資訊抽取 (network extraction)、4) 網路資訊正規化 (network normalization)、5) 映射 (mapping)、6) 分析 (analysis)以及7) 視覺化 (visualization)。特別要說明的是「資料前處理」是處理原始資料的重複和錯誤、區分時段(time slicing)以及網路資料縮減等工作,是決定科學映射分析能否得到良好結果的重要步驟之一。「網路資訊抽取」則是從論文的書目資料裡建立分析項目之間的關連,包括共現(co-occurrence)、耦合(coupling)和直接連結(direct linkage)等關係。兩個分析項目的共現關係取決於它們是否共同出現在一組文件內以及共同出現的次數;文件間的耦合關係則建立於它們是否具有共同的項目以及其數量大小,作者及期刊間的耦合關係則由屬於他們的文件的共同項目聚集而成;直接連結則是文件與它們的參考文獻之間的引用關係。運用不同的分析項目以及不同的關係可以對科學研究領域進行各種面向的分析,例如以文件中共同出現的作者所抽取的共同作者關係建立的網絡可以分析科學研究領域的社會結構(social structure);對於由詞語在文件內的共現關係所建構的詞語共現網絡進行分析則可以得知領域的概念結構(conceptual structure)和所處理的主要概念;經由文獻引用所產生的共被引關係和書目耦合關係則可以用來分析科學研究領域的知識結構(intellectual structure)。透過上面對於科學映射分析流程的分析,可以知道一個科學映射分析工具最好能夠具備以下的特性:a) 包含多種不同的模組來處理科學映射工作流程中的各個步驟;b) 具備強大的消除重複模組;c) 能夠建構各種書目計量的大型網絡;d) 具有良好的視覺化技術;e)輸出結果應該包含書目計量的測量結果與指標。本研究所提出的SciMAT工具具備上述的各種特性。SciMAT包含三個重要的模組:知識庫(knowledge base)、工作流程的配置以及測量結果與映射圖的視覺化模組。SciMAT的知識庫模組提供分析者匯入各種書目來源的檢索結果,將文件的作者、關鍵詞、期刊和參考文獻等各種資料儲存於知識庫內。運用此知識庫提供的功能,分析者能夠進行編輯與前處理等改善資料品質的工作以獲得更好的分析結果。SciMAT的工作流程配置模組循序漸進地設定分析的時間區段、分析的項目單位和關係、使用資料的次數閾值、進行資料正規化的相似性測量方式、叢集方式以及網絡分析、成效分析、時間分析和歷時性分析等相關參數。視覺化模組可以針對每個分析時段(period)提供詳細的網路圖、策略圖表以及相關的書目計量測量結果,也能夠提供代表研究主題(theme)的叢集在不同時段的演進情形等歷時性(longitudinal)的圖表

The general workflow in a science mapping analysis has different steps (Börner et al., 2003; Cobo et al., 2011b) (see Figure 1): data retrieval, data preprocessing, network extraction, network normalization, mapping, analysis, and visualization. At the end of this process, the analyst has to interpret and obtain conclusions from the results.

Usually, the data retrieved from the bibliographic sources contain errors, so a preprocessing process must be applied first. In fact, the preprocessing step is one of the most important to obtain good results in science mapping analysis. Different preprocessing processes can be applied to the raw data, such as detecting duplicate and misspelled items, time slicing, data reduction, and network reduction (for more information, see Cobo et al., 2011b).

A co-occurrence relation is established between two units (authors, terms, or references) when they appear together in a set of documents; that is, when they co-occur throughout the corpus.

A coupling relation is established between two documents when they have a set of units (authors, terms, or references) in common. Furthermore, the coupling can be established using a higher level unit of aggregation, such as authors or journals. That is, a coupling between two authors or journals can be established by counting the units shared by their documents (using the author’s or journal’s oeuvres).

Finally, a direct linkage establishes a relation between documents and references, particularly a citation relation.

In addition, different aspects of a research field can be analyzed depending on the units of analysis used and the kind of relation selected (Cobo et al., 2011b).

For example, using the authors, a coauthor or coauthorship analysis can be performed to study the social structure of a scientific field (Gänzel, 2001; Peters & van Raan, 1991).

Using terms or words, a co-word (Callon, Courtial, Turner, & Bauin, 1983) analysis can be performed to show the conceptual structure and the main concepts dealt with by a field.

Cocitation (Small, 1973) and bibliographic coupling (Kessler, 1963) are used to analyze the intellectual structure of a scientific research field.

We therefore think it would be desirable to develop a science mapping software tool that satisfies the following requirements: (a) it should incorporate modules to carry out all the steps of the science mapping workflow, (b) it should present a powerful de-duplicating module, (c) it should be able to build a large variety of bibliometric networks, (d) it should be designed with good visualization techniques, and (e) it should enrich the output with bibliometric measures.

SciMAT generates a knowledge base from a set of scientific documents where the relations of the different entities related to each document (authors, keywords, journal, references, etc.) are stored. This structure helps the analyst to edit and preprocess the knowledge base to improve the quality of the data and, consequently, obtain better results in the science mapping analysis.

Taking into account the GUI, there are three important modules: (a) a module dedicated to the management of the knowledge base and its entities, (b) a module (wizard) responsible for configuring the science mapping analysis, and (c) a module to visualize the generated results and maps. These modules allow the analyst to carry out the different steps of the science mapping workflow.

Regarding its functionalities, the module to manage the knowledge base is responsible for building the knowledge base, importing the raw data from different bibliographical sources, and cleaning and fixing the possible errors in the entities. It can be considered as a first stage in the preprocessing step.

As shown, the workflow is divided into four main stages: (a) to build the data set, (b) to create and normalize the network, (c) to apply a cluster algorithm to get the map, and (d) to perform a set of analyses. These stages and their respective steps are described below:
1. Build the data set: At this stage, the user can configure the periods of time used in the analysis (select the periods), the aspects that he or she wants to analyze (select the unit of analysis:  the conceptual (using terms or words), social (using authors), and intellectual (using references) aspects), and the portion of the data that has to be used (to filter the data using a minimum frequency as a threshold).
2. Create and normalize the network: At this stage, the network is built using co-occurrence or coupling relations or, indeed, aggregating coupling. Then, the network is filtered to keep only the most representative items. Finally, a normalization process is performed using a similarity measure (association strength (Coulter et al., 1998; van Eck &Waltman, 2007), Equivalence Index (Callon et al., 1991), Inclusion Index, Jaccard Index (Peters & van Raan, 1993), and Salton’s cosine (Salton & McGill, 1983).
3. Apply a clustering algorithm to get the map and its associated clusters or subnetworks: At this stage, the clustering algorithm used to build the map has to be selected. Different clustering methods are available in SciMAT, such as the Simple Centers Algorithm (Cobo et al., 2011a; Coulter et al., 1998), Single-linkage (Small & Sweeney, 1985), and variants such as Complete-linkage, Average-linkage, and Sum-linkage.
4. Apply a set of analyses: The final step of the wizard consists of selecting the analyses to be performed on the generated map.
(a) Network analysis: By default, SciMAT adds Callon’s density and centrality (Callon et al., 1991; Cobo et al., 2011a) as network measures to each detected cluster in each selected period. Callon’s centrality measures the degree of interaction of a network with other networks, and it can be understood as the external cohesion of the network. ... Callon’s density measures the internal strength of the network, and it can be understood as the internal cohesion of the network. ... These measures are useful to categorize the detected clusters of a given period in a strategic diagram (Cobo et al., 2011a).
(b) Performance analysis: SciMAT is able to assess the output according to several performance and quality measures. To do that, it incorporates into each cluster a set of documents using a document mapper function and then calculates the performance based on quantitative and qualitative measures (using citation-based measures, number of documents, etc.).
(c) Temporal analysis or longitudinal analysis: This allows the user to discover the conceptual, social, or intellectual evolution of the field. SciMAT is able to build an evolution map to detect the evolution areas (Cobo et al., 2011a) and an overlapping items graph (Price & Gürsey, 1975; Small, 1977) across the periods analyzed. Furthermore, SciMAT allows the user to choose different measures to calculate the weight of the “evolution nexus” (Cobo et al., 2011a) between the items of two consecutive periods, such as association strength (Coulter et al., 1998; van Eck & Waltman, 2007), Equivalence Index (Callon et al., 1991), Inclusion Index, Jaccard’s Index (Peters & van Raan, 1993), and Salton’s cosine (Salton & McGill, 1983).

At the end of all the steps in the wizard, the map would be built using the selected configuration. Then, the results would be saved to a file, and the visualization module loaded. The visualization module has two views: Longitudinal and Period.

The Period view (see Figure 12) shows detailed information for each period, its strategic diagram, and for each cluster, the bibliometric measures, the network, and their associated nodes.

Finally, in the Longitudinal view the overlapping map and evolution map are shown. This view helps us to detect the evolution of the clusters throughout the different periods, and study the transient and new items of each period and the items shared by two consecutive periods.

Taking into account quantitative measures such as the number of documents associated with each theme (cluster), we can discover where the fuzzy community has been employing a great effort (e.g., H-INFINITY-CONTROL, FUZZY-CONTROL, T-NORM, etc.). Similarly, taking into account the qualitative measure, we could identify the themes with a greater impact; that is, the themes that have been highly cited.

Combining the units of analysis and the bibliographic relations among them, SciMAT can extract 20 kinds of bibliographic networks, including the common bibliographic networks used in the literature, such as coauthor (Gänzel, 2001; Peters & van Raan, 1991), bibliographic coupling (Kessler, 1963), journal bibliographic coupling (Small & Koenig, 1977), author bibliographic coupling (Zhao & Strotmann, 2008), cocitation (Small, 1973), journal cocitation (McCain, 1991), author cocitation (White & Griffith, 1981), and co-word (Callon et al., 1983).