顯示具有 pathfinder scaling/minimum spanning tree 標籤的文章。 顯示所有文章
顯示具有 pathfinder scaling/minimum spanning tree 標籤的文章。 顯示所有文章

2015年4月15日 星期三

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V. (2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

Moya-Anegón, F. de, Vargas-Quesada, B., Chinchilla-Rodríguez, Z., Corera-Álvarez, E., Munoz-Fernández, F.J., & Herrero-Solana, V.(2007). Visualizing the marrow of science. Journal of the American Society for Information Science and Technology, 58(14), 2167–2179.

由於一般認為將領域之間的關係表示為圖形,通過考慮這些關係的可能性能夠提供許多資訊,不論對新進人員或專家皆有助於理解與分析,因此對這方面方法與工具的需求逐漸提高。過去的研究大多以期刊為分析單位,產生所有科學研究領域的科學映射圖。例如Leydesdorff (2004a, 2004b)使用雙重連結成分(biconnected components)的圖形分析演算法,將JCR 2001的科學研究進行分類。Boyack, Klavans, and Börner (2005)則應用了8種不同的期刊相似性測量7121種SCI和SSCI期刊,並採用VxOrd產生科學映射圖。Samoylenko, Chao, Liu, and Chen (2006)建構科學期刊的最小生成樹(minimum spanning trees),他們使用的資料是SCI 1994到2001的資料。本研究提出一個將ISI (Institute of Scientific Information)類別繪製成科學映射圖的方法,這個方法利用根據類別間的共被引資訊建構類別間的連結,以尋徑網路(PathfinderNetwork)縮減不重要的連結,然後以Kamada-Kawai方法決定節點在圖上的布局(layout),最後利用因素分析(factor analysis)進行結構確認。本研究和先前的研究都是針對類別利用共被引資訊呈現科學映射圖。以類別為分析單位在代表上足夠明確,並且比起較小的單位,這種方式對非專家使用者(nonexpert user)較具有資訊且使用者友善。Moya-Anegón et al. (2004)針對西班牙科學研究領域的視覺化,Moya-Anegón et al. (2005)則進一步利用科學映射圖比較英國、法國和西班牙三個國家的科學研究領域。本研究依循Börner, Chen, and Boyack (2003)提出的知識領域映射流程。使用的資料為7585種ISI期刊,ISI的類別共有219個,但扣除多學科科學後(Multidisciplinary Sciences),採用的類別共218個。利用共被引計算期刊相似性的方式為

Cc(ij)為期刊i和期刊j共被引次數,c(i)和c(j)則分別是期刊i和期刊j被引用次數。然後以尋徑網路和Kamada-Kawai方法繪製網路圖,經過尋徑網路處理後,有較多連結的節點具有較重要的地位。而尋徑網路是一種以型態為主的方法,與以群集為主的因素分析彼此間可以互補,因素分析可以識別、界定與定名科學映射圖上呈現的主題區域,而尋徑網路則負責讓使主題區域更加明顯,將類別分組成束,並顯示連接不同顯著類別的路徑,以及總體的型態結構。。最後總計共分析出35個因素,通過陡坡考驗(scree test)則有16個。科學映射圖上的類別可以分為三個群集:醫學與地球科學、基礎與實驗科學以及社會科學。

This study proposes a new methodology that allows for the generation of scientograms of major scientific domains, constructed on the basis of cocitation of Institute of Scientific Information categories, and pruned using PathfinderNetwork, with a layout determined by algorithms of the spring-embedder type (Kamada–Kawai), then corroborated structurally by factor analysis.

We present the complete scientogram of the world for the Year 2002.

This need arises from the general conviction that an image or graphic representation of a domain favors and facilitates its comprehension and analysis, regardless of who is on the receiving end of the depiction and whether a newcomer or an expert.

Science maps can be very useful for navigating around in scientific literature and for the representation of its spatial relations (Garfield, 1986). They are optimal means of representing the spatial distribution of the areas of research while also offering additional information through the possibility of contemplating these relationships (Small & Garfield, 1985).

From a general viewpoint, science maps reflect the relationships between and among disciplines; but the positioning of their tags clues us into semantic connections while also serving as an index to comprehend why certain nodes or fields are connected with others.

Moreover, these large-scale maps of science show which special fields are most productively involved in research—providing a glimpse of changes in the panorama—and which particular individuals, publications, institutions, regions, or countries are the most prominent ones (Garfield, 1994).

It is a tool in that it allows the generation of maps, and a method in that it facilitates the analysis of domains, by showing the structure and relations of the inherent elements represented. In a nutshell, scientography is a holistic tool for expressing the discourse of the scientific community it aspires to represent, reflecting the intellectual consensus of researchers on the basis of their own citations of scientific literature.

In Moya-Anegón et al. (2004), we ventured forth with a historic evolution of scientific maps from their origin to the present, and proposed ISI-JCR category cocitation for the representation of major scientific domains. Its utility was demonstrated by a visualization of the scientific domain of geographical Spain for the Year 2000.

Since then, other works related with the visualization of great scientific domains have appeared; however, all use journals as the unit of analysis, with the exception of a study based on the cocitation of categories (Moya-Anegón et al., 2005), comparatively focusing on three geographic domains (England, France, and Spain).

In contrast, Leydesdorff (2004a, 2004b) classified world science using the graph-analytical algorithm of biconnected components in combination with JCR 2001.

Boyack, Klavans, and Börner (2005) applied eight alternative measures of journal similarity to a dataset of 7,121 journals covering over 1 million documents in the combined Science Citation and Social Science Citation Indexes, to show the first global map of science using the force-directed graph layout tool VxOrd.

Samoylenko Chao, Liu, and Chen (2006) proposed an approach through the construction of minimum spanning trees of scientific journals, using the Science Citation Index from 1994 to 2001.

In processing and depicting the scientific structure of great domains, we further developed a methodology that follows the flow of knowledge domains and their mapping as proposed by Börner, Chen, and Boyack (2003).

Because ISI assigns each journal to one or more subject categories, to designate a subject matter (i.e., ISI category) for each document, we also downloaded the Journal Citation Report (JCR; Thomson Corporation, 2005a), in both its Science and Social Sciences editions, for 2002.

The downloaded records were exported to a relational database that reflects the structured information of the documents. This new repository contained nearly 1 million (N = 901,493) source documents: articles, biographical items, book reviews, corrections, editorial materials, letters, meeting abstracts, news items, and reviews that had been published in 7,585 ISI journals (N = 5,876 + 1,709). These were classified in a total of 219 categories, altogether citing 25,682,754 published documents.

As informational units, they are, in themselves, sufficiently explicit to be used in the representation of all disciplines that make up science in general. These categories, in combination with the adequate techniques for the reduction of space and the representation of the information to construct scientograms of science or of major scientific domains, prove much more informative and user friendly for quick comprehension and handling by nonexpert users than those obtained by the cocitation of smaller units of cocitation.

For these reasons, we used the 219 categories of the JCR 2002 as units of measure, with the exception of “Multidisciplinary Sciences.” ... The maximum number of categories with which we worked, then, was 218.

In light of our previous experience (Moya-Anegón et al., 2004, 2005), we use cocitation as the similarity measure to quantify the relationship existing between each one of the JCR categories.

Therefore, after a number of trials, we arrived at the conclusion that using tools of Network Analysis, the best visualizations are those obtained through raw data cocitation as the unit of measure. Yet, it also was necessary to reduce the number of coincident cocitations to enhance pruning algorithm yield. Therefore, to those raw data values we added the standardized cocitation value. In this way, we could work with raw data cocitation while also differentiating the similarity values between categories with equal cocitation frequencies. The key was a simple modification of the equation for the standardization of the degree of citation proposed by Salton and Bergmark:




where CM is cocitation measure, Cc is cocitation frequency, c is citation, and i and j are categories.

Over the history of the visualization of scientific information, very different techniques have been used to reduce n-dimensional space. Either alone or in conjunction with others, the most common are multidimensional scaling, clustering, factor analysis, self-organizing maps, and PathfinderNetworks (PFNET).

In our opinion, PFNET with pruning parameters r = ∞, and q = n − 1 is the prime option for eliminating less significant relationships while preserving and highlighting the most essential ones, and capturing the underlying intellectual structure in a economical way.

Although PFNET has been used in the fields of Bibliometrics, Informetrics, and Scientometrics since 1990 (Fowler & Dearhold, 1990), its introduction in citation was due to the hand of Chen (1998, 1999), who introduced a new form of organizing, visualizing, and accessing information. The end effect is the pruning of all paths except those with the single highest (or tied highest) cocitation counts between categories (White, 2001).

The spring embedder type is most widely used in the area of documentation, and specifically in domain visualization. Spring embedders begin by assigning coordinates to the nodes in such a way that the final graph will be pleasing to the eye (Eades, 1984). Two major extensions to the algorithm proposed by Eades (1984) have been developed by Kamada and Kawai (1989) and Fruchterman and Reingold (1991).

While Brandenburg, Himsolt, and Rohrer (1995) did not detect any single predominating algorithm, most of the scientific community goes with the Kamada–Kawai algorithm. The reasons upheld are its behavior in the case of local minima, its capacity to minimize differences with respect to theoretical distances in the entire graph, good computation times, and the fact that it subsumes multidimensional scaling when the technique of Kruskal and Wish (1978) is applied.

We can effortlessly see which are the most important nodes in terms of the number of their connections and, in turn, which points act as intermediaries with other lines, as hubs or forking points.

Whereas factor analysis is a clustering-oriented procedure, PFNET is topology oriented. Yet, they are extremely valuable as complements in the detection of the structure of a scientific domain.

Thus, factor analysis is responsible for identifying, delimiting, and denominating the great thematic areas reflected in the scientogram.

Meanwhile, PFNET is in charge of making the subject areas more visible, grouping their categories into bunches, and showing the paths that connect the different prominent categories, and finally, the overall topology of the domain.

Factor analysis identifies 35 factors in the cocitation matrix of 218 × 218 categories of world science 2002. Through the scree test we extracted 16, which we tagged using the previously explained method; these accumulate 70.2% of the variance (Table 1)

The number of categories included in at least one factor is 195. Twenty-three were not included in any factor (Table 2), and 25 belonged to two factors simultaneously (Table 5).

That is, a category or thematic area occupying a central position in the scientogram will have a more general or universal nature in the domain as a consequence of the number of sources it shares with the rest, contributing more to scientific development than those with a less central position.

The more peripheral the situation of a category or subject area, the more exclusive its nature, and the fewer the sources it will appear to share with other categories; accordingly, the lesser its contribution to the development of knowledge through scientific publications.

An intermediary position favors the interconnection of other categories or thematic areas. 

This broad interpretation of our scientograms not only explains the patterns of cocitation that characterize a domain but also foments an intuitive way for specialists and nonexperts to arrive at a practical explanation of the workings of PFNET (Chen & Carr, 1999).

From a macrostructural point of view, we can distinguish three major zones.

In the center is what we could call Medical and Earth Sciences, consisting of Biomedicine, Psychology, Etiology, Animal Biology & Ecology, Health Care & Service, Orthopedics, Earth & Space Science, and Agriculture & Soil Sciences.

To the right, we can see some other basic and experimental sciences: Materials Sciences & Physics, Applied; Engineering; Computer Science & Telecommunications; Nuclear Physics & Particles & Fields; and Chemistry.

To the left is the neighborhood of the social sciences, with Applied Mathematics, Business, Law, and Economy, and Humanities.

On one hand, it offers domain analysts the possibility of seeing the most essential connections between categories of given domain.

On the other hand, it allows us to see how these categories are grouped in major thematic areas, and how they are interrelated in a logical order of explicit sequences.

2014年2月28日 星期五

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. Journal of American Society for Information Science and Technology, 57(3), 359-377.

Chen, C. (2006). CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature.  Journal of American Society for Information Science and Technology, 57(3), 359-377.

information visualization

本研究提出一個整合研究專業(specialty)的研究前沿(research front)以及其引用的知識基礎(intellectual base)的視覺化介面。本論文定義研究前沿為研究專業上一組急遽出現的概念(concepts)與研究議題(research issues);研究前沿的知識基礎則是包含這些概念與研究議題的論文引用或者共同被引用的論文。在針對某一個專業進行其研究前沿與知識基礎進行視覺化時,首先蒐集專業相關的論文,從這些論文抽取代表研究前沿的詞語,並以論文所引用或共被引的論文做為專業的知識基礎,建立分別代表研究前沿的詞語和知識基礎的論文的二方網路(bipartite networks)以同時呈現研究前沿的相關概念與研究議題以及知識基礎的論文。在建立起來的網路上透過詞語和論文形成的叢集可以發現重要的研究前沿和知識基礎,藉由詞語呈現叢集的概念與研究議題更能有效地表達研究前沿的意涵,並且如果加上論文的發表時間來分析,可以從急遽出現在較多論文的相關詞語找出發展中的研究前沿。此外,對於網路進行中介中心性(centrality of betweenness)分析可以發現研究前沿間具有樞紐地位的論文,並且透過Pathfinder演算法可以發現論文間的主要關連。
A specialty is conceptualized and visualized as a time-variant duality between two fundamental concepts in information science: research fronts and intellectual bases.
A research front is defined as an emergent and transient grouping of concepts and underlying research issues.
The intellectual base of a research front is its citation and co-citation footprint in scientific literature— an evolving network of scientific publications cited by
research-front concepts.
The concept of a research front was originally introduced by Price (1965) to characterize the transient nature of a research field. Price observed what he called the immediacy factor: There seems to be a tendency for scientists to cite the most recently published articles. In a given field, a research front refers to the body of articles that scientists actively cite.
A specialty can be conceptualized as a time-variant mapping from its research front to its intellectual base.
Typical questions regarding a research front may include:
How did it get started? What is the state of the art? What are the critical paths in its evolution?
To address such questions, we need to detect and analyze emerging trends and abrupt changes associated with a research front over time. We also need to identify the focus of a research front at a particular time in the context of its intellectual base, to reveal significant intellectual turning points as a research front evolves, and to discover the interconnections between different research fronts.
Braam, Moed, and Raan (1991) defined a specialty as “focused attention by a number of scientific researchers to a set of related research problems and concepts” (p. 252). They studied the continuity and stability of a specialty in terms of the similarity between co-citation clusters across consecutive years. The similarity between two co-citation clusters is determined by comparing aggregated word profiles of the clusters.
In part, this is because we define a research front differently to emphasize emerging trends and abrupt changes as the defining features of a research front. A research front is the domain of a time-variant mapping, and its intellectual base is the co-domain of the mapping.
Griffith et al. (1974) found that between-cluster co-citation links tend to be weaker than within-cluster co-citation links. ... To understand how specialties and different thematic trends interact with each other, it is essential to study the nature of long-range, between-cluster links and understand why articles in different specialties were connected.
Labeling clusters is concerned with the clarity and interpretability of co-citation clusters. The standard approach relies on word profiles derived from articles citing a cluster of co-cited articles. ... Word-profile approaches have drawbacks. First, word profiles may not converge to a focused message. Analysts and users will make a substantial amount of sense-making efforts to synthesize a diverse range of word profiles. Second, cluster labels based on aggregating word profiles tend to be too broad to be useful. In practice, many users would be interested in not only the most commonly used terms but also terms that can lead to profound changes. Terms associated with an emerging trend could be overshadowed by a broader and more persistent theme.
In CiteSpace II, a current research front is identified based on such burst terms extracted from titles, abstracts, descriptors, and identifiers of bibliographic records. These terms are subsequently used as labels of clusters in heterogeneous networks of terms and articles.
CiteSpace II makes it easier for users to identify pivotal points. In addition to inspecting salient visual attributes, the user easily can see nodes with high betweenness centrality (Freeman, 1979).
The procedure of using CiteSpace II is described in the following steps, 
(1) Identify a knowledge domain using the broadest possible term.
(2) Data collection
(3) Extract research front terms: CiteSpace II first collects n-grams, or terms, from titles, abstracts, descriptors, and identifiers of citing articles in a dataset. The present study used single words or phrases of up to four words. ... Research-front terms are determined by the sharp growth rate of their frequencies.
(4) Time slicing
(5) Threshold selection
(6) Pruning and merging: Pathfinder network scaling is the default option in CiteSpace II for network pruning (Chen, 2004; Schvaneveldt, 1990).
(7) Layout
(8) Visual inspection
(9) Verify pivotal points
We demonstrate the new features of CiteSpace with case studies of two research fields: mass-extinction research (1981–2003) and terrorism research (1990–2003).
Mass-extinction research (1981–2003).
The input data for CiteSpace II were retrieved from citation index databases via the Web of Science based on a topic search for articles published between 1981 and 2003 on mass extinction. The scope of the search included four topic fields in each bibliographic record: title, abstract, descriptors, and identifiers. The search was limited to articles in English only.
The resultant dataset contains a total of 771 records.
A total of 333 research-front terms were detected from the four topic fields of these records.
Terrorism research (1990–2003).
The terrorism research (1990–2003) dataset consists of 1,776 records resulted from a topic search on terrorism in the Web of Science.
A total of 1,108 research-front terms were found.
The fully integrated representation of research fronts and intellectual bases in the same network visualization has three practical advantages.
First, using surged topical terms rather than the most frequently occurring title words is particularly suitable for detecting emerging trends and abrupt changes. In visualized networks, research-front terms are explicitly linked to intellectual-base articles. This design presents a compact representation of the duality between a research front and its intellectual base.
Second, research-front terms naturally lend themselves to be used as labels of specialties.
Third, it overcomes a common drawback of word-profile-based labeling approaches. Aggregated word profiles may not converge to an intrinsic focus. Terms selected based on sudden increased popularity measures are particularly suitable to characterize a current research front.
The Pathfinder algorithm extracts the most salient patterns from a network, but it does not scale well. CiteSpace II implements a concurrent version of the algorithm. The concurrent Pathfinder algorithm has substantially optimized the network scaling module, although it still took 6,000 seconds to process 14 networks and merge them into a 1,704-node network.
In conclusion, the new features introduced to CiteSpaceII for detecting and visualizing emerging trends and abrupt changes in a field of research have produced promising and encouraging results. The major findings are that
• the surge of interest is an informative indicator for a new research front;
• using heterogeneous networks of terms and articles provides a comprehensive representation of the dynamics of a specialty;
• research-front terms are informative cluster labels;
• citation tree-ring visualizations are visually appealing and semantically interpretable;
• betweenness centrality metrics identify semantically valid pivotal points.

2014年2月27日 星期四

Chen, C. and Morris, S. (2003). Visualizing evolving networks: minimum spanning trees versus pathfinder networks. In IEEE Symposium on Information Visualization 2003, Oct. 19-21, 2003, Seattle, Washington, USA, 67-74.

Chen, C. and Morris, S. (2003). Visualizing evolving networks: minimum spanning trees versus pathfinder networks. In  IEEE Symposium on Information Visualization 2003, Oct. 19-21, 2003, Seattle, Washington, USA, 67-74.

information visualization

本研究從網路的型態(topological)與動態(dynamical)兩方面比較Minimal Spanning Tree(MST)和Pathfinder Network(PFNet)兩種連結縮減的方法在共被引網路(co-citation network)的應用。連結縮減處理的目的在於使共被引網路能夠更清楚地呈現出重要的連結(論文或作者的共被引關係)與節點(論文或作者),因此一方面需要保留原本網路的型態,但另一方面當以論文引用的時間順序呈現時,能夠從網路的變化了解學科專業的演進。本研究發現經過MST處理的共被引網路會保留連結程度(degree)較高節點周圍的連結,所以可以維持原本網路的結構,但因為某些路徑上的重要連結被移除,因此不足以表達學科專業的演進;然而從PFNet所產生的網路不僅可以對應到學科專業的主題,同時在以共被引的時間呈現時也能夠發現到主題內以及主題之間的發展情形。
We compare the visualizations of co-citation networks of scientific publications derived by two widely known link reduction algorithms, namely minimum spanning trees (MSTs) and Pathfinder networks (PFNETs).
Two criteria are derived for assessing visualizations of evolving networks in terms of topological properties and dynamical properties.
The results suggest that although high-degree nodes dominate the structure of MST models, such structures can be inadequate in depicting the essence of how the network evolves because MST removes potentially significant links from high-order shortest paths. In contrast, PFNET models clearly demonstrate their superiority in maintaining the cohesiveness of some of the most pivotal paths, which in turn make the growth animation more predictable and interpretable.
The shortage of comprehensive examinations of the evolution of citation networks is due to various reasons, including the lack of an overarching framework that accommodates underlying theories and system functionalities across relevant disciplines, the lack of integrated network analysis and visualization tools, the lack of widely accessible longitudinal citation network data, and the lack of tools that specifically facilitate the analysis of network evolution.
A common problem with visualizing a complex network is that a large number of links may prevent users from recognizing salient structural patterns.
In fact, an MST is a special case of a Pathfinder network because a Pathfinder network is the set union of all the possible MSTs derived from a network [Schvaneveldt 1990].
In order to achieve a network of high clarity and legibility, it is necessary to impose the so-called triangular inequality throughout the network. While this requirement leads to the simplest representation of the essence of an underlying proximity network, this is at a considerable computational cost. Additionally, as the size of the original network increases, the algorithm requires a considerable amount of memory to run.
The most widely known graph drawing techniques include force-directed graph drawing algorithms and spring-embedder algorithms [Eades 1984]. ... These algorithms, however, face some challenges in terms of efficiency, especially in terms of scalability, which is closely related to the clarity of a visualized network.
A commonly used strategy to reduce clutter is to reduce the number of links. There are several ways to achieve this goal. Three popular ones are analyzed below.
The first option is imposing a link weight threshold and only include links with weights above the threshold [Zizi and Beaudouin-Lafon 1994]. ... However, it does not take the intrinsic structure of the underlying network into account, so the transformed network may not preserve the essence of the original network.
The second option is extracting a minimum spanning tree (MST) from a network of N vertices and reducing the number of links to N – 1.
The third option is imposing constraints on paths and excluding links that do not satisfy the constraints, for instance, as in Pathfinder network scaling [Schvaneveldt 1990]. ... The topology of a PFNET is determined by two parameters q and r and the corresponding network is denoted as PFNET(r, q). The q-parameter specifies the maximum length of a path subject to the triangular inequality test. The r-parameter is the Minkowski metric used to compute the distance of a path. The most concise PFNET for visualization is PFNET (q = N–1, r = inf) [Chen 2002; Chen and Paul 2001; Schvaneveldt 1990].
Most network growth models draw upon the rich-get-richer notion and cumulative advantage. As a result, if the degree of a node indicates its “richness,” a node with a higher degree will have a better chance to receive the next new link than a node with lower degree. In a citation network, this means that a highly cited article is more likely to be cited again than a less frequently cited article. This type of growing mechanism is known as preferential attachment.
It appears to be particularly problematic to identify significant topological and dynamical patterns in such visualization models because of the high density of the underlying network.
An et al. [2001] suggested that the evolution of citation networks could be useful in predicting research trends and in studying a scientific community’s life span.
Two criteria are derived based on the above analysis for qualitatively evaluating network visualization.
The first criterion for selecting a preferable topological structure of a visualized network is the presence of hubs, or stars, in derived networks. ... A star pattern indicates the star node carries the most information, processes the highest cue validity and the most differentiated from one another. ... Existing studies appear to suggest that co-citation counts are likely to form such star patterns in both MST and PFNET.
Criterion II requires that the changes of topological properties over time must preserve the integrity of emergent trends or patterns. Visualizing network evolution should not merely inform users of changes of individual nodes and links; rather, it is essential to inform users how an intrinsically cohesive structure changes locally and globally in organically.
In this study, research fronts were identified by agglomerative clustering using only papers that had at least five bibliographic coupling counts with some other paper in the dataset. Similarity calculation was based on Salton's cosine coefficient [Salton 1989] applied to bibliographic coupling counts. The titles for each research front were derived manually by exploring titles of papers within each research front for common themes.
Base reference clusters were formed by agglomerative clustering using only references that had been cited 10 or more times. Similarity calculation was based on Salton's cosine coefficient applied to co-citation counts. For each base reference cluster, labels were found by using the label of the research front that contained the most citations to references in the cluster.
A map of the references in the pathfinder network was produced identifying each reference by its base reference cluster membership, which allowed labeling of sections of the pathfinder network based on base cluster labels.
In general, due to the arbitrary choice inherited from the MST algorithms, one cannot guarantee the uniqueness of an MST. As a result, an MST may not preserve all the necessary links for representing the growth of a co-citation network. If this is the case, then important diffusion patterns may be distorted or inadequately represented by the extracted MST model.
The 516-node PFNET (q = N – 1, r = inf) is shown in Figure 3. The two parameters q and r were chosen to ensure that the extracted PFNET has the least number of links.
The animated PFNET visualization model demonstrated that nodes with similar colors often emerged simultaneously and formed local structures. And these local structures were reinforced by the timely emergence of salient co-citation links. The growth process can be represented by the dynamics shown in such local structures. Features such as continuity, predictability, and local cohesiveness in the PFNET indicated that the second criterion was met.

Chen, C. (1997). Structuring and visualizing the WWW with Generalized Similarity Analysis. Proceedings of the 8th ACM Conference on Hypertext (Hypertext '97), 177-186.

Chen, C. (1997). Structuring and visualizing the WWW with Generalized Similarity Analysis. Proceedings of the 8th ACM Conference on Hypertext (Hypertext '97), 177-186. Retrieved August 27, 2012, from http://delivery.acm.org/10.1145/270000/267456/p177-chen.pdf?ip=211.76.242.1&acc=ACTIVE%20SERVICE&CFID=108526370&CFTOKEN=44188441&__acm__=1346053732_db4b154c3eaabf4b1429f52a8e30ba0c
vis_paper

本論文以PathFinder方法提供網頁(或網站)視覺化的呈現,並且根據網頁彼此間的超文件連結(hypetext linkage)、內容相似度(content similarity)和瀏覽樣式(browsing patterns)來衡量它們的接近度(proximity)。具體而言,本研究利用網頁間的連結數目比率、向量空間模式(vector space model)以及網頁間的狀態轉移機率(state transition probability)來估算彼此間的接近度。以網頁做為圖形(graph)上的節點(vertices),網頁間估測的接近度做為節點間連結線的強度,然後再藉由Pathfinder方法在保留網絡的主要型態下,去除不必要的連結,作者認為Pathfinder比多維尺度法(MDS)能夠更精確地表現圖形在區域間的關係(local relationship)。
This paper describes a generic approach to structuring and visualizing a hypertext-based information space on the WWW. This approach, called Generalised Similarity Analysis (GSA), provides a unifying framework for extracting structural patterns from a range of proximity data concerning three fundamental relationships in hypertext, namely, hypertext linkage, content similarity and browsing patterns.
Pathfinder networks are used as a natural vehicle for structuring and visualizing the rich structure of an information space by highlighting salient relationships in proximity data.
Georgia Institute of Technology’s WWW User Surveys [17] shows that 69. 1% of users regarded the delay in downloading Web pages as a major problem and 34.5% of users identified the difficulty of finding an existing page. In particular, 14.3% of the users reported the difficulty of visualizing where they have been and where they can go and 6.5% identified the classic hypertext problem — lost in hyperspace. The memory overload remains a problem when navigating the WWW.
Ideally, spatial relationships in visualization should be determined by some psychological judgments of proximity, such as similarity, dissimilarity and relatedness.
Pirolli, Pitkow and Rae’s study [18] and HyPursuit [20] are two notable examples of taking into account hypertext linkage, content similarity and usage information on the WWW.
In HyPursuit, document similarity by linkage is defined as a linear combination of three components: direct linkage, ancestor and descendant inheritance.
Pirolli, Pitkow and Rao [18] developed a model which characterises documents on the WWW by various attributes associated with these documents, such as the number of incoming and outgoing hyperlinks of a document, how frequently the document was downloaded from the hosting WWW server and content similarities between the document and its children.
Sequential patterns of browsing indicate, to some extent, document relatedness perceived by users. For example, the number of users who followed a hyperlink connecting two documents in the past were used in [18] to indicate the degree of relatedness between the two documents.
Furnas’ fisheye views model is based on a “degree of interest” (DOI) function which assigns a value to each node in accordance with the degree to which a user would be interested in seeing that node [14, 12]. ... A fisheye view can be generated with a threshold so that only nodes with sufficient DOI are displayed in the view. ... By choosing a different API function, one can produce a fisheye view which emphasizes a particular type of structural patterns[ 12]. For example, the number of times that a node has been visited can be used to define a user-centred fisheye view, in which popular nodes will be highlighted for easy access.
In this paper, we focus on extracting underlying relationships in a hypertext information space and representing resultant patterns for structuring and visualizing the information space. Existing techniques such as fisheye views can be subsequently incorporated into such systems with improved spatial configuration mechanisms.
This definition also takes into account the overall connectivity of the document Di, which can be related to the ROC metric defined in [2].
In this study, we use the well-known tf x idf model, term frequency times inverse document frequency, to build term vectors. ... The document similarity is computed as follows based on corresponding vectors.
We have applied a state transition approach to extracting behavioral patterns of users with a hypertext system [6]. The dynamics of a browsing process can be captured by state transition probabilities. Transition probabilities can be used to indicate document similarity in the nature of browsing.
Pathfinder provides a more accurate representation of local relationships than techniques such as multidimensional scaling (MDS)[10]. Pathfinder has been applied to a number of human-computer interaction problems [10].
The topology of a PFNET is determined by two parameters q and r and the corresponding network is denoted as PFNET(r,q). The q-parameter constrains the scope of minimum-cost paths to be considered. The r-parameter defines the Minkowski metric used for computing the distance of a path.
When a PFNET satisfies the following 3 conditions, the distance of a path is the same as the weight of the path:
1. The distance from a document to itself is zero.
2. The proximity matrix for the documents is symmetric; thus the distance is independent of direction.
3. The triangle inequality is satisfied for all paths with up to q links. If q is set to the total number of nodes less one, then the triangle inequality is universally satisfied over the entire network.
The number of links in a network can be reduced by increasing the value of parameter r or q. The distance between nodes in a network is the length of the minimum-length path connecting the nodes; such a path is known as the geodesic connecting the nodes. A minimum-cost network (MCN), PFNET(r=INF, q=n- 1), has the least number of links.
The major advantage of Pathfinder networks is that salient relationships among documents are extracted by patterns associated with minimum-cost paths. This type of information filtering improves the clarity and quality of the information produced by information visualization systems based on spring models. Users are able to see how documents are related to each other.
GSA has some distinct features. 
(1) GSA emphasizes that users can substantially benefit from explicit, graphical representations of salient relationships in hypertext systems, and these graphical representations should be incorporated into user interfaces so as to reduce cognitive burdens on users in browsing.
(2) Each component model in GSA can be used independently for extracting structures of a particular type so that users may contrast patterns in distinct characteristics. In contrast, related work such as [ 18] combines various features into a monolithic feature vector, Consequently, the resulting inter-document relationship is a combined effect of a range of factors. Users may not be able to assess how documents are related along a specific dimension.
(3) GSA focuses on relationships that are particularly essential for hypertext systems and these relationships are preserved in resulting network representations. Many existing information visualization techniques are based on storage information such as file-size and last modification time, and often use hierarchical structures as the basis of visualization. Differences between the two approaches should be evaluated by further empirical studies.
For example, a Pathfinder network becomes increasingly cluttered as the number of documents in the underlying information space increases. There are several possible ways to deal with this issue. One is to use existing display techniques such as fisheye views, which provide adequate access to specific local information as well as contextual structure.
Similar documents are naturally placed near to each other in the space. Users can gain a birds-eye view of the global structure by moving up to a higher view point in the sky and have a close look by moving down to a view point closer to the target document.

Chen, C. M., & Paul, R. J. (2001). Visualizing a knowledge domain's intellectual structure. Computer, 34(3), 65-71.

Chen, C. M., & Paul, R. J. (2001). Visualizing a knowledge domain's intellectual structure. Computer, 34(3), 65-71.
vis_paper
本論文進行ACA(author citation analysis)的研究,以IEEE Computer Graphics and Applications上發表論文的作者為分析對象,選擇353位被引用5次以上的作者,利用他們之間的共被引資訊建立網路圖,結果共有28,638條連結線。經過尋徑網路尺度(pathfinder network scaling)的處理,保留下355條比較重要的連結線。為了發現電腦圖學與應用的專長(specialties),本論文借鏡於White and McCain(1998)的研究,利用PCA(principal component analysis)方法對共被引資料進行因素分析(factor analysis),結果共得到60個專長,5個較大的專長共可以解釋39%的變異數,而這5個專長分別是Rendering and ray tracing、Computer vision、Geometric modeling and computer-aided design、Volume rendering和Modeling nature。同時也在網路圖上呈現被歸類為這5個專長的作者,來觀察他們在網路圖上的分布情形。
ACA, a special type of citation analysis, focuses on intellectual connections between authors as reflected through the scientific literature. The author co-citation relationship links two authors by how often other authors reference their work together. Author co-citation patterns provide the basis for constructing an alternative view to a knowledge structure.
Pathfinder uses a filtering criterion known as the triangle inequality condition to determine whether to remove or retain each link in the original network. Triangle inequality requires that the length of a path connecting two points in the network should not be longer than the length of other alternative paths connecting the two points, but go through extra intermediate points.
We began by studying author co-citation patterns found in IEEE Computer Graphics and Applications magazine for a period of 18 years. ...  Among them, we entered into the author co-citation analysis only the 353 authors who received more than five citations in CG&A. Although this snapshot derives from a limited viewpoint—the literature of  computer graphics certainly stretches beyond the scope of CG&A— intellectual groupings of these 353 authors provide the basis for visualizing the computer graphics knowledge domain. ... The original author co-citation network contains as many as 28,638 links, which constitutes 46 percent of all possible links, excluding self-citations. Because this many links would clutter visualizations, we applied Pathfinder network scaling to reduce their number to 355.
We enhanced the network by coloring it according to the results generated using principal component analysis (PCA). PCA identified 60 specialties in computer graphics. The largest (rendering and ray tracing) and second-largest (computer vision) accounted for 13 percent and 11 percent of the variance, respectively. The five largest specialties accounted for 39 percent of the variance. Remaining specialties are relatively small.
Factor 1: Rendering and ray tracing.
Factor 2: Computer vision.
Factor 3: Geometric modeling and computer-aided design.
Factor 4: Volume rendering.
Factor 5: Modeling nature.
The knowledge landscape visualizes intellectual structures. A virtual landscape like this provides an intuitive gateway for users to access the scientific literature. Researchers new to a field can gain a useful overview by using the knowledge landscape to establish their own mental model of the field and track the development of their own domain.

Chen, C., & Carr, L. (1999). Trailblazing the literature of hypertext: Author co-citation analysis (1989-1998). Proceedings of the 10th ACM Conference on Hypertext (Hypertext '99), 51-60.

Chen, C., & Carr, L. (1999). Trailblazing the literature of hypertext: Author co-citation analysis (1989-1998). Proceedings of the 10th ACM Conference on Hypertext (Hypertext '99), 51-60.
vis_paper
本論文以9屆(1987-1998)的ACM Hypertext 學術研討會會議論文為研究資料,運用作者共被引分析(author co-citation analysis, ACA)、Pearson相關係數分析(Pearson’s correlation coefficients)、因素分析(factor analysis)等技術,探討超文件處理與應用學術領域的研究專長(specialties),並利用尋徑網路尺度(Pathfinder network scaling)將研究專長分析所產生的結果進行視覺化。在這個研究裡,共分析367位引用次數較多的作者之間的共被引現象,結果共產生39個因素,這些因素共解釋了87.8%的變異數。若以前四個因素而言,則解釋了52.1%。從因素內的作者來命名,前四項超文件處理與應用學術領域的研究專長分別是經典(Classics)、資訊檢索(Information retrieval)、圖形使用者介面與資訊視覺化(Graphical user interfaces and information visualisation)以及連結與連結機制(Links and linking mechanisms)。
The ultimate goal of our work is to realise the vision of making the best use of an interrelated information space and building one’s own threads of association. As one step in this direction, we explore a new paradigm of structuring and visualising a domain-specific information space.
In this study, we choose the field of hypertext as the subject domain and map the literature of hypertext based on the ACM Hypertext conference proceedings (1987-1998).
The idea of mapping the tracks of science is explained by Garfield in [8]. The aim of such work is to identify research front specialties in a field of study. A specialty is characterised by its influence on the development of a given field. One can tell a specialty by the number of citations that it receives.
In 1981, Institute for Science Information (ISI) published ISI Atlas of Science in biochemistry and molecular biology [10]. The Atlas was constructed based on co-citation index associated with publications in the field over a limited period of one year. 102 distinct clusters of articles were identified, which were called research front specialties, in order to give researchers a snapshot of significant research activities in biochemistry and molecular biology.
White and McCain [17] used author co-citation analysis to map the field of information science. ... Their study also included a factor analysis, in which major specialties were identified. One of the most remarkable findings is that the field of information science consists of two major specialties with litter overlap between their memberships: experimental retrieval and citation analysis.
In a series of studies, we have been investigating the role of Pathfinder network scaling techniques in reducing the excessive number of links and extracting the most salient structures from a range of proximity data [3]. One problem we repeatedly encountered is an interpretation problem: users found hard to make sense the nature of links selected by Pathfinder. ... A simple and easy-to-understand method is needed to explain the structure of a Pathfinder network, especially when the nodes are high dimensional in nature.
Following [17], the raw co-citation counts were transformed into Pearson’s correlation coefficients using the factor analysis. These correlation coefficients were used to measure the proximity between authors’ co-citation profiles. ... In the factor analysis, principal component analysis with varimax rotation was used to extract factors. The default criterion, eigenvalues greater than one, was specified to determine the number of factors extracted. ... Pearson correlation matrices were submitted to the GSA environment for processing, especially including Pathfinder network scaling and VRML-scene modelling.
Thirty-nine factors were extracted from the 367 x 367 author co-citation data set. These factors explain 87.8% of the variance. In particular, the top four factors alone explain 52.1% of the variance.
Factor 1: Classics.
Factor 2: Information retrieval.
Factor 3: Graphical user interfaces and information visualisation.
Factor 4: Links and linking mechanisms.
Pathfinder networks can provide more accurate information about local structures than multidimensional scaling maps [13]. We found that the provision of explicit links in our maps made it easier to interpret interrelationships among different data points.
Furthermore, author co-citation maps provide a means of identifying research fronts, i.e. specialties in the field, and a visual aid of interpreting the results of factor analysis.

2014年2月8日 星期六

Chen, C., McCain, K., White, H., & Lin, X. (2002). Mapping Scientometrics (1981–2001). Proceedings of the American Society for Information Science and Technology, 39(1), 25-34.

Chen, C., McCain, K., White, H., & Lin, X. (2002). Mapping Scientometrics (1981–2001). Proceedings of the American Society for Information Science and Technology, 39(1), 25-34.

科學映射圖(science mapping)是整合資訊視覺化(information visualization)和科學計量學(scientometrics)的研究,藉由圖形呈現揭露科學文獻的結構與相關的專業(specialties),科學映射圖的最基礎技術為詞語的共現分析和共被引分析,分別提供獨特的科學研究前沿結構洞察力,Braam, Moed, & Raan (1991a, 1991b)的研究發現結合這兩種技術能夠讓出版品的認知內容(cognitive content of publications)產生更為清楚的圖像。

科學計量學是測量科學或技術進展的研究 (Garfield, 1979b)。傳統上科學計量學有相當強烈的應用導向,針對科學或技術的輸入與輸出發展出許多測量方法與指標,許多知識工作者以這些測量方法與指標為工具應用於各種不同的研究:例如可以針對國家、地區和研究機構的研發能力進行政策與計畫的評估,或是對於研究領域的知識結構進行領域分析。van Raan (1997)和Persson (2000)都是以Scientometrics期刊論文做為研究資料的研究。van Raan (1997)分析科學計量學的最佳狀態(the state of the art)以及對它的應用導向傳統進行描述,van Raan建議科學計量學需要和知識發現(knowledge discovery)與資料探勘(data mining)整合來獲得明顯的效益。Persson (2000)以1978到1999年,44卷,1062篇論文資料進行分析,找出最常被引用的出版品,並且產生期刊共被引、國家間的直接引用連結、作者間的共被引以及直接引用等圖形,表現各種不同的結構。

本研究以1981到2001年間的Scientometrics期刊論文為研究資料,選擇被引用次數達五次以上的參考文獻,共計403筆文獻,根據這些文獻的共被引資訊,繪製網路圖做為科學映射圖的基本圖形,並以論文的引用速率產生動畫的效果。本研究首先以文獻的共被引次數計算Pearson 相關係數(Pearson's correlation coefficients)產生共被引矩陣(co-citation matrix)。並且利用主成分分析(principal component analysis)對共被引矩陣進行因素分析,以分析出的因素代表領域的專業。同時也利用共被引矩陣產生網路圖,經過尋徑網路縮放(pathfinder network scaling)保留較重要的共被引連結,以簡化圖形的複雜性。最後以VRML(virtual reality modeling language)呈現圖形,並且以動畫呈現文獻的被引用率增長情形。本研究共計找出25個因素,較大的三個因素所對應的專業分別命名為科學研究中的引用(citations in science studies)、全球與國家的科學表現(world and national science performance)、研究產出的評估(evaluation research outputs)。

The design of the visualization model adapts a virtual landscape metaphor with document cocitation networks as the base map and annual citation rates as the thematic overlay. The growth of citation rates is presented through an animation sequence of the landscape model.

Science mapping aims to reveal structures of scientific literature and underlying specialties using graphical representations. ... Co-word analysis (Callon, Law, & Rip, 1986) and co-citation analysis (Small, 1973) are among the most fundamental techniques for science mapping. ... Each offers a unique perspective on the structure of scientific frontiers. Researchers have found that a combination of co-word and co-citation analysis could lead to a clearer picture of the cognitive content of publications (Braam, Moed, & Raan, 1991a, 1991b).

As an integral part of our long-term research, our investigation emphasizes an interdisciplinary synergy that may involve fields of study such as information visualization and scientometrics.

Can we provide domain analysts, science performance evaluators, researchers, students, and other knowledge workers something tangible and meaningful that they can readily incorporate it into their work process? Can we improve the way we learn about a new subject matter, the way we explore a knowledge domain, and the way we trace the history and evolution of a specialty? And ultimately, can we augment our ability to judge the significance of scientific work more efficiently and more accurately?

The present study is based on articles published in Scientometrics between 1981 and 2001, drawn from the Web of Science.

Scientometrics is “the study of the measurement of scientific and technological progress” (Garfield, 1979b). Its origin is in the quantitative study of science policy research, or the science of science, which focuses on a wide variety of quantitative measurements, or indicators, of science at large.

Input indicators include the amount of research grants awarded to institutions and the number of people receiving scientific degrees; output indicators include the number of scientific articles published, the number of citations to each article, and the number of patents granted.

Science policy and program evaluation studies have used such indicators to measure the scientific strength of various countries, regions, or research institutions.

Domain analysts have used such indicators to describe the intellectual structure of a knowledge domain.

Scientometric research has a strong application-oriented tradition (Garfield, 1979b; Raan, 1997).

Garfield (Garfield, 1979b) identified several publications appeared in the 1970s and contributed to the development of scientometrics, namely, the first Science Indicators published by the National Science Board in 1972 (Board, 1977), the Evaluative Bibliometrics: The Use of Publication and Citation Analysis in the Evaluation of Scientific Activity by Francis Narin and Computer Horizons, Inc. (CHI) in 1976 (Narin, 1976), which has been regarded as a good review source for anyone interested in scientometrics (Garfield, 1979b).

Derek Price’s 1965 article ‘Network of Scientific Papers’ (Price, 1965) has been also regarded as a key event in the development of the field of scientometrics.

Michael Moravcslk (1977) presented a review of scientometric literature (Moravcslk, 1977).

Anthony van Raan (1997) analyzed the state of the art of scientometrics and characterized its application-oriented tradition. He envisaged that scientometrics could benefit significantly from a greater integration with knowledge discovery and data mining.

Loet Leydesdorff (2001) identified some challenges of scientometrics and suggested that: “the state of the art of science studies is ‘preparadigmatic:’ it is an interdisciplinary area integrated only at the level of its subject matter, and an applicational area for various contributing disciplines.”

A directly related study of Scientometrics was done by Olle Persson (2000). He retrieved 1,062 articles published in the journal from volume 1 to volume 44 between 1978 and 1999. Top-10 most cited publications include (Garfield, 1979a; Schubert, Glanzel, & Braun, 1989; Small, Sweeney, & Greenlee, 1985). He generated several maps to show a variety of structures, including journal co-citation, direct citation links among countries, shared citations among authors, and direct citations among authors.

This study is based on bibliographc data retrieved from the Web of Science. The data contain all types of documents published in Scientometrics between 1981 and 2001. ... Each article must be cited for 5 times or more in order to be included in this integrated analysis. This threshold resulted in a total of 403 articles.

In this study we have adapted an integrated procedure of citation analysis and information visualization, including Pathfinder network scaling, Principal Component Analysis (PCA), and visual-spatial models rendered in Virtual Reality Modeling Language (VRML).

The cocitation strength is computed as Pearson’s correlation coefficients to form a co-citation matrix. ... The co-citation matrix forms the basis of a base map, a terminology commonly used in cartography.

Factor analysis, namely PCA, is subsequently applied to the co-citation matrix in order to produce a thematic overlay. The purpose of such a thematic overlay is to highlight the density distribution of various specialties. Factor loadings are used to color code each publication in the thematic overlay.

We simplify the cocitation matrix using Pathfinder network scaling, which retains the strongest co-citation links with reference to the so-called triangle inequality condition (Chen, 1997, 1998; Schvaneveldt, 1990).

Finally, the growth of citation rates is animated across the entire Pathfinder network to facilitate the identification of trends. The visualization-animation model is made available in VRML 2.0 for easy access on the Internet.

PCA identified 25 factors from the 403 by 403 co-citation matrix. In theory, each factor should correspond to a specialty. ... The large number of factors reflects the diversity of scientometrics.

In our analysis, we focus on the three largest factors of significant specialties of the field.
Specialty 1: Citations in Science Studies.
Specialty 2: World and national science performance.
Specialty 3: Evaluation research outputs.

2013年12月19日 星期四

White, H. D., Lin, X., Buzydlowski, J. W. and Chen, C. (2004). User-controlled mapping of significant literatures. Proceedings of the National Academy of Science of the United States of America, 101, 5297-5302.

White, H. D., Lin, X., Buzydlowski, J. W. and Chen, C. (2004). User-controlled mapping of significant literatures. Proceedings of the National Academy of Science of the United States of America, 101, 5297-5302.

本研究使用尋徑者網路(pathfinder networks, PFNET)和自組織映射圖(self-organizing maps, SOM)兩種維度縮減(dimension reduction)技術做為PNAS期刊論文檢所的圖形化介面,輸入一個詞語或一位作者,產生這個查詢與其相關的24個詞語或24位作者的圖形。以Gene Frequency與這個主題的重要作者Slatkin做為查詢所產生的PFNET與SOM,提供Slatkin檢視,他認為產生的圖形很容易解釋。本研究並說明與比較了PFNET和SOM作為資訊視覺化界面的特點。

information visualization

Our data are the contents of PNAS for 1971–2002, as described by medical subject headings from the National Library of Medicine (NLM) and by citation indexing from the Institute for Scientific Information (ISI).
SOMs show frequently co-occurring terms as nodes that are spatially close. PFNETs show them as nodes with explicit ties. The two kinds of maps will be exemplified here with medical subject headings (MeSH) and cocited authors in a specialty of genetics.
In their extensive review, Borner et al. (5) emphasize that ‘‘painting a big picture’’ is a main goal in domain mapping. This may lead to a strategy of mapping very large co-occurrence matrices in their entirety. Indeed, system designers have made many significant developments in software for such global portrayals of literatures, e.g., THEMESCAPE and VXINSIGHT render literatures as landscapes; GALAXIES and STARRYNIGHT render them as astral bodies (10–12).
In global mapping, system designers present the user with a preformed view, often in 3D, of some sizeable literature. Within the panel of visualization, landscapes invite flyovers; star-fields or other constructs invite flythroughs. In the former, peaks representing major accretions of documents on some subject are likely to exert a powerful pull on the user; in the latter, document points coded as important, e.g., by differences in shape, size, or color, exert a similar pull.
Essentially, the user is engaged in old-fashioned browsing, as of book titles in library stacks, but system designers may minimize or even eliminate labeling of objects in the map because labels clutter precious screen space and block the metaphorical presentation (see examples in ref. 12).
The user explores the view by ‘‘visiting’’ or ‘‘homing in on’’ objects of interest, rather as in video games, but typically cannot remap the literature in pursuit of some new interest because a new map takes hours of computer time to create.
Ours, however, is an alternative way of visualizing knowledge domains, the localized mapping. Perhaps the chief difference is that the localized approach relinquishes scope to increase the user’s control of the mapping process.
... our localized system of mapping more closely resembles online searching. The user starts the process by entering a single term at a web interface. This is consistent with the way most people search the web (13) and is intended to minimize cognitive demands on users.  The system responds to the entry (or ‘‘seed’’) term by forming a list of the terms that co-occur with it, ranked high to low by frequency. The seed term and its 24 next-highest neighbors are then exhibited as a PFNET or a SOM, which the user can switch between.
If the indexing terms used in the mapping are indeed controlled by a formal thesaurus, our SOMs and PFNETs provide an alternative: they display the top listings in what is sometimes called a term’s associative thesaurus (2).
A map of cocited authors is, in effect, an associative thesaurus of authors linked by conjoint use of their works. Again, these linkages may permit useful retrievals that are not otherwise possible (1).
PFNETs and SOMs are dimension-reduction techniques that have been used to visualize the structure of literatures for more
than a decade.
In the context of the movement joining bibliometrics with document retrieval (2, 5, 10), PFNETs have been described by Fowler and colleagues (15–17), McGreevy (18), and Chen (19, 20). Analogous accounts of SOMs have been done by Lin et al. (21), Roussinov and Chen (22), and Chen et al. (23).
The number of links in a PFNET is controlled by two parameters, r and q.
The parameter r, which determines how path weights are computed, is lucidly explained by Fowler et al. (17): ‘‘Path weight, r, is computed according to the Minkowski r-metric. It is the rth root of the sum of each distance raised to the rth power for all links in a path between two nodes. Although the r-metric is continuously variable, simple interpretations exist only for r =1 (path weight is the sum of the link weights in the path), r=2 (path weight is the Euclidean distance), and r=infinity (path weight equals the maximum link weight in the path). One advantage of r=infinity is that one need only assume that the original distance estimates have ordinal properties. Another advantage is that the link structure will be preserved for anymonotonic transformation of the data.’’
The parameter q sets the range within which all paths of length q will be examined in the test of the triangle inequality (24) and removed if they violate it. The larger the value of q, the more extensive the triangle inequality constraint; therefore, links are more likely on a path that violates the rule. If q is one less than the number of nodes, then all of the potential violators are under scrutiny.
The more frequently co-occurring terms, which presumably have greater mutual relevance, occupy more proximate regions on the map. SOMs are designed to render not just the highest co-occurrence counts between terms, but rather relatively high co-occurrences across groups of terms.
They are a softer-focus kind of mapping than PFNETs, but they, too, suggest specific combinations of terms on which the user might want to base retrievals.
This process of self-organization (also known as unsupervised learning) runs over many iterative cycles. In each iteration, the images of term pairs that are strongly related in the high-dimensional space will be moved closer on the lower-dimensional space until stability is reached.
A row from the cooccurrence matrix ‘‘is randomly selected and compared to every output node to determine a winner. Weights of the winning output nodes then are updated so that the next time this input node is presented, this output node will likely be selected again as the winner. In the meantime, nodes surrounding the winning node are similarly adjusted.
The number of iterations needed to train a SOM is often determined empirically (in our case, we optimize the number of training cycles to 2,500).
After the training, input vectors closest in the input space will map to the same regions in the output map. The regions are delineated by areas of nodes in which the elements with the highest value on the vectors are the same.’’
Adjacent areas reflect stronger relationships than nonadjacent areas. Terms in large areas are more influential than terms in small areas.
Slatkin found his own cocited author maps readily interpretable. He was acquainted with every name that appears in Fig. 2. In the PFNET (which he again preferred), he identified the main structural feature, the clusters around himself and Masatoshi Nei, as representing two slightly different subject areas. Both the Nei group and the Slatkin group, he said, have contributed to the literature on genetic flow and population structure, but the Slatkin group has contributed relatively more to the literature on microsatellites (short, repetitive sequences of DNA). Hence, the PFNET was picking up a division he found meaningful.
Interestingly, at the lower left the SOM conjoins Wright, Mayr, and Fisher, who represent the older, pioneering generation in statistical genetics. The SOM algorithm is able to bring this out solely on the basis of their overall cocitation profiles.
If PFNETs seem directive about term relationships, SOMs are merely suggestive. However, their greater ambiguity is perhaps a virtue.
Using AUTHORLINK, the forerunner of PNASLINK, Buzydlowski (9) found that SOMs outperformed PFNETs in capturing the mental models of 20 experts in selected fields of the humanities. ... The experts’ mental models were elicited by having them sort cards bearing authors’ names into intuitively meaningful piles. ... SOMs agreed with the card-sort data better than PFNETs. In the Plato trial, both SOMs and PFNETs were highly correlated with the pooled card-sort data (SOMs, r 0.97; PFNETs, r 0.78), but these correlations were significantly different at P 0.001. In the individual-author trials, a t test of mean agreement scores favored SOMs significantly at P<0.01. 

2013年5月15日 星期三

Quirina, A., Cordóna, O., Vargas-Quesadab, B., Moya-Anegón, F. (2010). Graph-based data mining: A new tool for the analysis and comparison of scientific domains represented as scientograms. Journal of Informetrics, 2010, 291-312.


Quirina, A.,  Cordóna, O., Vargas-Quesadab, B., Moya-Anegón, F. (2010). Graph-based data mining: A new tool for the analysis and comparison of scientific domains represented as scientograms. Journal of Informetrics, 2010, 291-312.

network analysis
本研究利用基於圖形的資料探勘(graph-based data mining)方法,嘗試從數個科學圖表(scientograms)中發現它們之間共同的次結構(substructures),藉以分析以下的三個問題:1)全國科學研究領域的逐年進展情形,2) 世界各國共同研究類別次結構(research categories substructures)的抽取以及3)不同國家間科學研究領域的比較。本研究首先對分析的73個國家每一年的科學研究進行研究類別的共被引分析,建構以每一個研究類別為節點、類別間的共被引相關程度為連線的網路圖做為本研究探討的科學圖表,每一條連線的權重計算方式為CM(ij) = Cc(ij) + Cc(ij)/sqrt(c(i)*c(j)),其中c(i)和c(j)分別是i與j兩個類別的被引次數,Cc(ij)則是i與j兩個類別的共被引次數。當某兩對類別的共被引次數不相等時,由於 Cc(ij)/sqrt(c(i)*c(j))的值介於0與1之間,此一權重計算方式對於共被引次數較大者可以得到較大的數值;當兩對類別的共被引次數相等時,兩個類別的被引次數接近於共被引次數的情況下,可以得到較大的數值。然後將建立起來的網路圖經過尋徑者網路(pathfinder networks)演算法進行維度縮減處理(dimensionality reduction)。進行處理時,將尋徑者網路演算法的參數r和q,分別將r和q設定為r =∞ , q = n − 1。對於任意兩個節點間的連結線,如果除了此一連結線,以這兩個節點為端點的任何的一條路徑,其上的某一個連結線的權重大於此連結線的權重,則此連結線的重要性較不顯著(significant),便於網路圖上刪除此連結線。網路圖可以利用Kamada–Kawai等佈局演算法(layout algorithm),將節點與其之間由線連結所表現的關係呈現為視覺化圖像。本研究採用Subdue演算法(Cook & Holder, 1994, 2000)抽取各網路圖共同的次結構,Subdue演算法的運算原理是最小描述長度(minimum description length) (Rissanen, 1989)。將網路圖上的節點、連結線以及節點間的連結關係表示成位元串(bit string),位元串的長度總和便是網路圖的描述長度。假設各網路圖的集合為G,描述G的位元串的長度總和為I(G),並假定出現在多個網路圖中的各次結構為S,如果描述S的位元串的長度總和為I(S),將各網路圖上出現的此次結構以一節點取代的情況時,描述G的位元串的長度總和成為I(G|S)。最小描述長度的原理便是發現各個次結構使得I(S)+I(G|S)的值最小,實際上以valueMDLi(S,G) = I(G) / (I(S) + I(G|S))求取最大值。在評估各次結構時,除了求得最小描述長度(也就是最大的valueMDLi),另外次結構的選擇以規模(size)較大、支持度(support)較大者也是比較具有代表性的次結構。規模以網路圖與次結構中包含的節點和連結線總數來計算,評估方式為valuesize(S,G) = Size(G) / (Size(S) + Size(G|S))。支持度則是次結構被包含於多少網路圖內來計算,評估方式為valuesupport (S,G) = #graphs in G including S / card(G),card是網路圖的數目。如果考慮負圖表的情形,也就是希望次結構出現於負圖表愈少愈好,則描述長度的評估方式成為valueMDLi(S,Gp,Gn) = (I(Gp) + I(Gn)) / (I(S) + I(Gp|S) + I(Gn) − I(Gn|S)),其中Gp和Gn分別代表正圖表與負圖表;規模的評估方式成為valuesize(S,Gp,Gn) = (Size(Gp) + Size(Gn)) / (Size(S) + Size(Gp|S) + Size(Gn) − Size(Gn|S));支持度則成為valuesupport (S,Gp,Gn) = (#Gp graphs including S + #Gn graphs not including S) /  (card(Gp) + card(Gn))。以下分別說明本研究探討的三個問題的實驗方式與結果:
1)全國科學研究領域的逐年進展情形
本研究以年度為單位,要分析的年度所產生的網路圖為正圖表,先前的年度為負圖表,發現顯著包含於正圖表內的次結構。
2) 世界各國共同研究類別次結構的抽取
以世界各國為單位,所有的國家所產生的網路圖都視為是正圖表,發現顯著包含於正圖表內的次結構。
3)不同國家間科學研究領域的比較
以世界各國為單位,要分析的國家所產生的網路圖為正圖表,其餘國家的網路圖為負圖表,發現顯著包含於正圖表內的次結構。
In this paper, we aim to show that graph-based data mining tools are useful to deal with scientogram analysis. Subdue, the first algorithm proposed in the graph mining area, has been chosen for this purpose. This algorithm has been customized to deal with three different scientogram analysis tasks regarding the evolution of a scientific domain over time, the extraction of the common research categories substructures in the world, and the comparison of scientific domains between different countries.
The visualization of scientific information has long been used to uncover and divulge the essence and structure of science (Börner & Scharnhorst, 2009; Chen, 1999a, 2004).
Yet despite its ripe age, information display is still in an adolescent stage of evolution in the context of its application to scientific domain analysis.
There is a large number of information visualization techniques which have been developed over the last decade within this area (Chen, 1999b; Lucio-Arias & Leydesdorff, 2008; Moya-Anegón et al., 2007, 2005; Small & Garfield, 1985), but none of them has been designed to support the exploration of large datasets. Besides, all the latter approaches require a large amount of expertise from the user, which reduces the chances to automate the analysis procedure. Nevertheless, it is clear that information visualization and visual data mining (Keim, 2002) can provide the theoretical and practical backgrounds to deal with scientific information analysis.
The generation of a big picture is something implicit in the process of visualizing scientific information. In an attempt to sum what has taken place to date up, we can say that nowadays there are two proposals for tracking down the big picture. On the one hand, one can adopt the traditional units of analysis (authors, documents, and journals) and, through their grouping, identify scientific disciplines following a bottom-up process (Boyack & Klavans, 2008; Klavans & Boyack, 2006; Small & Sweeney, 1985; Small, Sweeney, & Greenlee, 1985). On the other hand, the alternative uses the categories of the documents to the same end, and shows the scientific structure from them in a top-down manner (Moya-Anegón et al., 2004).
The former proposal (bottom-up process) presents all the pros of its fine-grained character, but it runs into difficulties in representing the totality of the panorama on a single plane and in tagging the disciplines.
That is, it (top-down process) is relatively simple to represent the scientific structure of a domain on a single plane by means of a maximum of 300 categories and their interrelation, avoiding tagging problems. However, this implies the acceptance of a classification of science in predefined categories, never transparent and always subjective, as well as the fact that documents are classified by the journals in which they are published and not by their content (coarse-grained character).
Current scientogram analysis techniques (Boyack, Börner, & Klavans, 2009; Chen et al., 2009; Klavans & Boyack, 2006; Leydesdorff & Rafols, 2009; Moya-Anegón et al., 2007) aim to provide a fine, detailed, tight view of a scientogram. To do so, they are based on performing a low-level analysis and comparison of the maps. Statistical techniques, computer algorithms, and macrostructure and microstructure techniques for the identification of thematic areas and scientific disciplines have already been used to analyze and compare scientograms (Boyack, Klavans, & Börner, 2005; Chen, 1999b; Lucio-Arias & Leydesdorff, 2008; Moya-Anegón et al., 2007; Wallace, Gingras, & Duhon, 2009).
However, this approach shows a main limitation: only a single or a very reduced set of maps can be analyzed or compared together. In fact, the field lacks an easy-to-use approach allowing the identification and the comparison of scientific structures within scientograms with a higher degree of automation.
Graph-based data mining (GBDM) (Cook & Holder, 2006; Holder & Cook, 2005; Washio & Motoda, 2003) involves the automatic extraction of novel and useful knowledge from a graph representation of data. By ‘novel’ we mean that the knowledge retrieved is not directly encoded in the data but deeply masked in it (hence, it requires to be uncovered), and by ‘useful’ we mean that the discovered patterns have in general an interest for the domain expert ...
In fact, GBDM techniques have been applied for frequent substructure discovery and graph matching in a large number of domains including chemistry and applied biology (Borgelt & Berthold, 2002; Huan et al., 2004), classification of chemical compounds (Deshpande, Kuramochi, & Karypis, 2002), and unsupervised and supervised pattern learning (Cook & Holder, 2006), among many others.
Subdue (Cook & Holder, 1994, 2000) is a graph-based knowledge discovery system that finds structural, relational patterns in data representing entities and relationships. It aims to discover interesting and repetitive substructures in a structural database (DB). For this purpose, the minimum description length (MDL) principle (Rissanen, 1989) is used in order to compress the original DB into a hierarchical and shorter version.
In particular, we will describe how this algorithm can be customized to deal with three different scientogram analysis and comparison tasks regarding the evolution of a scientific domain over time, the extraction of the common research categories substructures in the world, and the comparison of scientific domains between different countries.
The generation of a scientogram following the top-down approach (Moya-Anegón et al., 2004) requires the sequential application of several techniques.
1. Units of analysis: The categories are the units of analysis and representation (Moya-Anegón et al., 2004; Vargas-Quesada & Moya-Anegón, 2007).
2. Unit of measure: ... a co-citation measure CM is computed for each pair of categories i and j as follows:
CM(ij) = Cc(ij) + Cc(ij)/sqrt(c(i)*c(j))
where Cc is the co-citation frequency and c is the citation frequency.
3. Dimensionality reduction: ...  the Pathfinder algorithm (Chen, 1998; Dearholt & Schvaneveldt, 1990) is applied to the co-citation matrix to prune the network. Due to the density of the data, and especially in the case of vast scientific domains with a high number of entities (categories in our case) in the network, Pathfinder is usually parameterized to r =∞ and q = n − 1. This is done in order to preserve and highlight the salient relationships between categories, and for capturing the essential underlying intellectual structure of a scientific domain.
4. Layout: The spring embedder family of methods is the most widely used in the area of Information Science. Spring embedders assign coordinates to the nodes in such a way that the final graph will be pleasing to the eye, and that the most important elements are located in the center of the representation (also called its backbone). Kamada–Kawai’s algorithm (Kamada & Kawai, 1989) is one of the most extended methods to perform this task. Starting from a circular position of the nodes, it generates networks with aesthetic criteria such as the maximum use of available space, the minimum number of crossed links, the forced separation of nodes, the build of balanced maps, etc.

The need of mining structural data to uncover objects or concepts that relates objects (i.e., subgraphs that represent associations of features) has increased in the past ten years, thus creating the area of GBDM (Holder & Cook, 2005; Washio & Motoda, 2003). Nowadays, GBDM has become a very active area and several techniques such as Subdue, the Apriori family of methods (Apriori-based GM (Inokuchi, Washio, & Motoda, 2000), Frequent Subgraph Discovery (Kuramochi & Karypis, 2001), JoinPath (Vanetik, Gudes, & Shimony, 2002), etc.), and the Frequent Pattern-growth family of methods (CloseGraph (Yan & Han, 2003), FFSM (Huan, Wang, & Prins, 2003), Gaston (Nijssen & Kok, 2004), gSpan (Yan & Han, 2002), MoFa/MoSS (Borgelt & Berthold, 2002), Spin (Huan, Wang, Prins, & Yang, 2004), etc.) have been proposed to deal with problems such as graph matching, graph visualization, frequent substructure discovery, conceptual clustering, and unsupervised and supervised pattern learning (Cook & Holder, 2006).
Among them, we can highlight Subdue (Cook & Holder, 1994, 2000), a graph-based knowledge discovery system that finds structural, relational patterns in data representing entities and relationships. ... It is able to develop graph shrinking as well as frequent substructure extraction and hierarchical conceptual clustering.
Subdue (Cook & Holder, 1994, 2000) is a method for discovering interesting and repetitive substructures in a structural DB. The algorithm uses the MDL principle (Rissanen, 1989) to discover frequent substructures in a DB, extract them and replace them by a single node in order to compress the DB. These extracted substructures represent structural concepts in the data.
The Subdue algorithm can be run several times in a sequence in order to extract meta-concepts from the previously simplified DB. After multiple Subdue runs on the DB, we can discover a hierarchical description of the structural regularities in the data (Jonyer, Cook, & Holder, 2001).
Subdue can also use background knowledge, such as domain-oriented expert knowledge, to be guided and to discover substructures for a particular domain goal.
Subdue uses a variant of beam search (Lowerre, 1976) in order to avoid exponential-sized queue: at each step, only BeamWidth new children from a given parent are explored (see line 14). Furthermore, only a maximum of MaxBest substructures having a maximal size of MaxSubSize are returned to the user, and the algorithm does not develop more than Limit iterations (see line 6). These parameters ensure that the running time of Subdue is polynomial and is actually constrained by the BeamWidth and the Limit parameters (Jonyer et al., 2001).
The evaluation of a substructure (see line 13) can be computed by the MDLmeasure (see Section 3.1.1), the Size-measure (see Section 3.1.2), or the Support-measure (see Section 3.1.3).
The MDL of a graph is the necessary number of bits for describing completely the graph. This number of bits is usually given by the value I(S), the number of bits required to encode the substructure S. I(S) is computed as the sum of the number of bits to encode the vertices of S, the number of bits to encode the edges of S, and the number of bits to encode the adjacency matrix describing the graph connectivity of S. Subdue looks for the substructure S minimizing I(S) + I(G|S), where G is the input graph, I(S) is the number of bits required to encode the uncovered substructure, and I(G|S) is the number of bits required to encode the graph obtained by compressing G with S, i.e., substituting each occurrence of S in G by a single node (Holder, Cook, & Djoko, 1994).
In the following, we renamed the MDLi measure (’i’ stands for inverse) as we are maximizing its value: Subdue considers a given substructure S is better than another one S if the MDLi measure valueMDLi(S,G) is higher than valueMDLi(S,G), where valueMDLi(S,G) is computed as follows: valueMDLi(S,G) = I(G) / (I(S) + I(G|S))
However, the alternative operation mode for Subdue considers two distinct sets, a positive set Gp and a negative set Gn, determined by the user. In this operation mode, the goal of Subdue is to find the largest substructures present in the maximum number of graphs in the positive set, which are not included in the negative set. The MDLi measure is thus computed as follows:
valueMDLi(S,Gp,Gn) = (I(Gp) + I(Gn)) / (I(S) + I(Gp|S) + I(Gn) − I(Gn|S))
The size of an object is not computed from the description length, but from an index based on either the number of nodes, the number of edges or, more usually, the sum of the both values. This measure is faster to compute but less consistent as it does not show the real benefit obtained after the compression of the DB.
valuesize(S,G) = Size(G) / (Size(S) + Size(G|S))  where, usually, Size(G) = #vertices(G) + #edges(G).
In the case of the second operation mode, in which we have a positive and a negative scientogram set, the Size measure is computed as follows:
valuesize(S,Gp,Gn) = (Size(Gp) + Size(Gn)) / (Size(S) + Size(Gp|S) + Size(Gn) − Size(Gn|S))
The last alternative measure is based on the support of substructure S and it is expressed as follows:
valuesupport (S,G) = #graphs in G including S / card(G) with card(G) being the cardinal (cardinality ?) of the set of graphs G composing the DB.
For the second operation mode, this evaluation measure is computed as the sum of the number of positive maps containing S and the number of negative maps not containing S, divided by the total number of maps. Its formulation is as follows:
valuesupport (S,Gp,Gn) = (#Gp graphs including S + #Gn graphs not including S) /  (card(Gp) + card(Gn))

We consider a substructure having a larger positive support and a smaller negative support as having a better quality. In the same way, substructures having a larger size are preferred over smaller ones as they are more specific.

Study of the evolution of the scientific domain of a specific country over time
An information science expert would be interested in knowing which substructures appear in the analyzed domain, at which time, how big they are, how many they are, where are they located, and so forth. This will allow him to perform at least two kind of studies. On the one hand, an in-deep analysis of the uncovered substructures themselves, which kind of categories are they linking, etc. On the other hand, global statistics about the size and the quantity of these substructures to respectively characterize the importance of the evolution of the domain and its dynamics.
Thus, the goal of the first analysis task is to present a framework for the study of the evolution of a scientific domain over time using Subdue. ... As we want to look for CRCSs which were appearing at a given time, we also need to pick two ranges of years, the negative range and the positive range. The negative range is usually a set of years from the past, in which these substructures (i.e., CRCSs) are not meant to exist. The positive range is usually a set of years dated after the negative range, in which the substructures are meant to be present.

Identification of the common research categories substructures in the world
The aim of the second scientogram analysis task is to uncover the CRCSs in the world by analyzing the scientograms of a large number of different countries. ... All the selected maps representing the scientific production of those countries for that given year will be viewed as positive examples, so the goal of Subdue will be to extract the substructures with the best support among all of them. Notice that, no negative examples are considered in this case. As the user will be specially interested on the extracted CRCSs to be as specific as possible, the MDLi measure will be again considered to extract both frequent and large substructures.

Comparison of the scientific domains of different countries
To do so, the scientogram of each country in a given set is compared against the remaining ones in that set, the current country viewed as a positive map and the others as negative maps. ... Note that this experiment could also be done using time periods larger than a single year, or more than one country in the positive set each time, thus allowing an expert to extract the substructures highlighting the possible similarities between these countries.