2013年12月19日 星期四

Schildt, H. A. and Mattsson, J. T. (2006). A dense network sub-grouping algorithm for co-citation analysis and its implementation in the software tool Sitkis. Scientometrics, 67, 143-163.

本研究建議一個密集網路次群集演算法(dense network sub-grouping algorithm)從共被引網路裡發現學術領域的主流研究主題。在以論文被其他論文的引用情形做為代表論文的特徵,利用Jaccard指標(Jaccard index)衡量論文間在共被引方面的相關性,建立網路圖。然後利用密集網路次群集演算法將網路圖分為幾個節點集合和無法歸入集合的節點。本研究並討論了建議的方法與共被引分析常用的叢集分析(cluster analysis)和多維縮放(MDS, multidimensional scaling),密集網路次群集演算法不須事先決定結果集合的數目,而且對可以歸入多個集合的論文(通常有較廣泛的共被引關係)有較佳的處理方式。本研究將密集網路次群集演算法應用於家庭事業(family business)的研究,同時也發展了Sitkis軟體將ISI的檢索結果轉換成可以做為常用於網路分析的軟體UCINET輸入的檔案格式。

information visualization

We propose an alternative algorithm, dense network sub-grouping, which identifies dense groups of co-cited references. We demonstrate the algorithm using a data set from the field of family business research and compare it to two alternative methods, multidimensional scaling and clustering.
The software identifies journal-, country- and university-specific citation patterns and co-citation groups, enabling the identification of “invisible colleges.”
Gmür’s recent article (2003) provides a review of such methodologies and suggests that clustering algorithms may be useful in identifying research streams or “invisible colleges” among scientists within a field.
Because pre-existing clustering algorithms were not designed for bibliometric analysis, several factors make them suboptimal for the task. First, optimizing clustering algorithms requires the number of clusters to be defined ex ante, whereas hierarchical clustering algorithms lack clear boundaries between clusters.
More important, clustering algorithms always assign each article into a cluster, with no residual category for items that resemble many disparate clusters. This may potentiall cause broadly cited articles to reside in one cluster over another based on very slight differences in citation patterns. The relative prominence of clusters, measured in terms of citation counts, tends to depend heavily on these “citation classics.” Coincidentally, many highly cited works are also cited relatively broadly. As a result, the apparent popularity of different approaches within a field may not be robust to small variations in citation patterns. To avoid this, very broadly cited articles should be excluded altogether from the clusters.
The dense network sub-grouping algorithm is based on iterative identification of the tightly coupled areas of co-citation networks.
The algorithm starts formation of a group at the most strongly connected dyad, and then iteratively adds dyads one at a time, ordered by highest average tie strength to existing group members. When average tie strength from pre-existing group members to all other nodes is below the given cut-off value, the algorithm terminates. The newly formed group is removed from the network, and the algorithm begins to search for the next group. Each group represents a cohesively cited body of literature.
The empirical part of this paper suggests that these groups tend to correspond either to a specific empirical research question or a theoretical approach.
A widely accepted tool for meta-analysis, bibliometrics has been applied in social science disciplines including: economics (CAHLIK, 2000; PIETERS & BAUMGARTNER, 2002), finance (BOROKHOVICH et al., 2000; CHUNG & COX, 1990; HOLLMAN et al., 1991; SCHWERT, 1993), strategic management (MARTINSONS et al., 2001; RAMOS-RODRIGUEZ & RUIZ-NAVARRO, 2004), entrepreneurship (BUSENITZ et al. 2003; DERY & TOULOUSE, 1996; RATNATUNGA & ROMANO, 1997), inter-organizational relationships (OLIVER & EBERS, 1998; SOBRERO & SCHRADER, 1998), organization studies (ÜSDIKEN & PASADEOS, 1995), marketing (PASADEOS et al., 1998), and research and development studies (TIJSSEN & VAN RAAN, 1994).
The algorithm groups together the works that are commonly cited together in scientific articles. All members of such a group would have a similar subject and readership. Since the most-cited prior works in a scientific field arguably represent its key intellectual roots, combining the works into groups based on co-citation coupling will provide a “map” of the field’s intellectual structure (GMÜR, 2003).
A new algorithm, dense sub-network grouping, was developed for the purposes of co-citation analysis. ... It begins by forming a group at the dyad that has the highest co-citation value and then iteratively adds nodes ordered by the highest average co-citation link to the existing members of the group, until the average link value is lower than a predetermined cut-off value chosen by the researcher. The resulting group is then removed from the network, and the algorithm proceeds from the beginning.
We advocate that co-citation networks be constructed using a normalized co-citations strength measure, the Jaccard index (SMALL & GREENLEE, 1980). The normalization is used in order to emphasize proximity between similar references that are cited less often than the most common references.
The distinct advantage of our algorithm is that it omits widely cited books that do not belong to any coherent stream of literature. Also, because we focus on coherent groups of references, our analysis allows us to include works that are less cited, but clearly part of a coherent stream.
Based on this sample of outlets for publication, we used the Institute of Scientific Information Social Sciences Citation Index (ISI SSCI) to systematically select all family business related articles published during the period 1986 – 2003. Acknowledging the definitional diversity of “family business research” (SHARMA, 2004), we searched for variations on the terms “family firm,” “family business,” “family ownership,” and “family control.” An initial set of 341 articles was obtained, which was then systematically reviewed to select out any articles not related to family business. Altogether 108 articles were included in the final sample. Subsequently, we selected all references that had been cited by at least four (3.7%) of these 108 articles.
A major shortcoming of the traditional clustering method was evident from the beginning of the analysis: because the resulting clusters must be evaluated manually, the number of works included in the analysis has to be kept manageably low. This reduces the precision of the analysis and may leave the cluster structure imperfect or skewed. ... One can manually construct about 8 to 12 relevant clusters, depending on the threshold at which different clusters are allowed to join together. However, the manual cluster formation is not a straightforward task. In addition, because of the relatively small number of works included, the clusters themselves are rather small, making the results less comprehensive.
Generally, interpreting the MDS results requires much speculation on the part of the researcher, and thus may not yield replicable results.
There are four common reference data incoherencies: (1) authors’ middle initials are used inconsistently; (2) journal and book names are spelled inconsistently in reference information (e.g. ADMIN SCI Q and ADM SCI Q represent the same publication); (3) different publication years (editions) of the same book appear as independent entries; and (4) multiple articles by the same author in the same journal in the same year appear as a single entry.

Chalmers, M. (1992). BEAD: Explorations in information visualisation. Proceedings of the 1992 ACM SIGIR Annual International Conference on Research and Development in Information Retrieval, 330-337.

Chalmers, M. (1992). BEAD: Explorations in information visualisation. Proceedings of the 1992 ACM SIGIR Annual International Conference on Research and Development in Information Retrieval, 330-337.
vis_paper
本論文說明Bead的運作原理。Bead是以圖形探索文件資訊的雛型系統,在這個系統中將文件視為是三維空間中的粒子(particles),類似像粒子間的作用力,使得相似的文件在空間中相對應的粒子移動到愈接近的距離,不相似的文件移動到較遠的位置。也就是以粒子在三維空間中的幾何距離(geometric distance)表示文件距離(document distance),以輸出結果的三維空間表達原先相當高維度的文件空間。在這個雛型系統中粒子的位置是由多維尺度法(MDS)來決定,以文件間的距離做為輸入產生每個相對應的粒子在三維空間中的位置。並且利用模擬退火(simulated annealing)演算法來避免產生區域最佳化,而有較佳的輸出結果。
Bead is a prototype system for the graphically-based exploration of information. In this system, articles in a bibliography are represented by particles in 3-space. By using physically-based modelling techniques to take advantage of fast methods for the approximation of potential fields, we represent the relationships between articles by their relative spatial positions. Inter-particle forces tend to make similar articles move closer to one another and dissimilar ones move apart. The result is a 3D scene which can be used to visualize patterns in the high-D information space.
As described in [Chatfield 80], MDS is “a term used to describe any procedure which starts with the ‘distances’ between a set of points (or individuals or objects) and finds a configuration of the points, preferably in a smaller number of dimensions, usually 2 or 3.” Such a configuration may then afford the graphically-based, exploratory styles of manipulation usually reserved for lower dimensional data.
Standard techniques for MDS usually involve an eigenvector analysis of an nxn matrix, where n is the number of points to be configured. This O(n3) analytic procedure generates the required configuration of points in a single step, but changes or additions to the set of points require re-execution of the entire analysis.
As mentioned by Chatfield & Collins with regard to a special case of MDS, ordinal scaling, one can treat the configuration task as an optimization problem. This has allowed the introduction of iterative numerical techniques such as the steepest descent method [Press 88]. ... One problem of steepest descent is that it may be unable to climb out of a local minimum. The points become stuck in a configuration which disallows progress towards a ‘better’ configuration.
Applying these principles to numerical optimization led to the Metropolis algorithm for simulated annealing [Metropolis 53]. This requires a way of describing possible configurations (e.g. the spatial position of points), a method of generating random changes in the configuration (e.g. perturbation of position), an objective function, analogous to energy, whose minimization is the goal of the procedure (e.g. the error between the actual and desired inter-point distances), and finally a control parameter analogous to temperature, in tandem with a schedule for its reduction. ... As the schedule progresses, random changes in configuration are successively generated. At each step the change in system energy is calculated. If the step lowers the energy then the change is accepted. If the step increases the energy then the change is accepted with a probability which depends on both the size of the energy step and the current temperature of the system. ‘Uphill’ steps are taken less often as temperature decreases.
The behaviour of a body of matter can be considered as the aggregate of the behaviours of its constituent particles. ... Potentially, all N particles can lead to N(N -1) pairwise interactions. This N-body problem is particularly well-known in computational physics where interactions due to gravitational or Coulombic fields are often of exactly this O(N2) complexity. ... Approximate solutions of lower orders of complexity have therefore been subjects of research activity in recent years. By using hierarchical structures for spatial subdivision such as k-d trees [Appel 85] and octrees [Barnes 86] relatively straightforward algorithms of O(N logN) complexity have been developed.
We associate documents with particles in space, and use two concepts to underpin this work: a cluster of such hybrid objects can be represented by a metaparticle, and when comparing particles, the difference between the actual ‘geometric’ distance and the desired ‘document’ distance can be used as the basis for a potential field. ... We use a model of a damped spring in order to generate forces of attraction and repulsion between particles. When particles are too close the spring pushes them apart, and when they are too far apart they are drawn towards each other. ... We use hierarchical spatial subdivision in order to speed up the calculation of the interactions between the full set of particles. We also use the hierarchy to form document clusters, using addition and normailization of term vectors.
We use a 3-D tree to hierarchically subdivide the space into which particles are placed. We maintain a metaparticle for each node in the tree, consisting of a particle position, a mass, and a term vector. We use document distances to metaparticles to recursively descend the tree in an effort to insert a new document near to similar documents. We maintain a maximum number of particles inside each leaf node — usually 10 — and split a leaf when that limit is reached. In splitting voxels we try to find an axis and coordinate with which to split the number of particles as evenly as possible.
Finally, we iteratively use fourth-order adaptive Runge-Kutta in applying either a simulated annealing step or a steepest descent step to the particle system. In this way we attempt to minimize the unbalanced force on each particle.In this way we try to find positions where the often-conflicting demands for proximity reach the best available compromise.
While looking at the maximum unbalanced force on any particle in the system is one useful measure of progress, we have also found it informative to look at the residual sum of squares. Analogous to the mechanical stress of the spring system and denoted here as S, this metric is built up by examining each pair of particles a and b. ... We assume that an ‘acceptable’ configuration has been reached when continued iteration produces roughly constant levels of stress and maximum force.

McCain, K. W. (1990). Mapping authors in intellectual space: a technical overview. Journal of the American Society for Information Science, 41(6), 433-443.

McCain, K. W. (1990). Mapping authors in intellectual space: a technical overview. Journal of the American Society for Information Science, 41(6), 433-443.

vis_paper

本論文說明作者共被引分析(author cocitation analysis, ACA)的進行步驟與相關技術,ACA的分析流程包括1)選取即將分析的作者集合、2)取得作者的共被引次數、3)建立原始的作者共被引矩陣、4)利用原始共被引矩陣計算相關係數,每一對作者之間以他們與其他作者共被引次數分布的相似程度作為他們之間的接近值、5)對接近值矩陣進行叢集分析(cluster analysis)、多維尺度分析(multi-dimensional scaling, MDS)和因素分析(factor analysis),產生視覺化圖形、6)解釋與驗證。
Within a given map, the proximity of points representing authors reflects their perceived similarity on some dimension. By examining the distribution of authors and author clusters within the two- or three-dimensional “intellectual space” of a mapped display, other aspects of structure can be described. Clusters of points can be identified with subject areas, research specialties, schools of thought, shared intellectual styles, or temporal or geographic ties. In a factor analysis, factor loadings may demonstrate the breadth or concentration of various authors’ scholarly contributions.
A common sequence of steps in author cocitation analysis is as follows,
1) Selection of the author set: One relatively objective way to identify potentially well-cited authors is to choose those who have many page references in a text, monograph, or collection of review articles.
2) Retrieval of cocited author counts
3) Compilation of  raw cocitation matrix
4) Conversion of the raw data matrix to a matrix of proximity values: The creation of a correlation matrix has at least two major advantages. First, for any given pair of authors, the correlation coefficient functions as a measure, not just of how often that pair of authors were cocited (the raw frequency count), but of how similar their “cocitation profiles” are. ... The correlation coefficient also removes differences in “scale” between authors who are highly cited and those who have similar profiles but are less frequently cited overall (Kerlinger, 1973). ... The correlations are defined as measures of similarity: the higher the positive correlation, the more similar two authors are in the perceptions of citers.
5) Approaches to multivariate analysis have been used to display the inter-author relationships in the similarities matrix:
a) In ACA, cluster analysis is used to group authors so as to provide insights into the intellectual organization of a given field. ... The two most popular approaches to cluster formation are called “hierarchical agglomerative” vs. “iterative partitioning” ... ACA research has tended to use the agglomerative clustering approach. The hierarchical agglomerative methods can use the correlation matrix as similarity measures among the authors. Authors are paired, an author is joined to an existing cluster, or two clusters are fused based on their similarity.
b) Multidimensional scaling (MDS) requires as input the same matrix of similarities or dissimilarities among objects as cluster analysis, and the two are often used together. MDS is a set of techniques used to create visual displays- maps -from proximity matrices, so that the underlying structure within a set of objects can be studied. In ACA, the major uses of multidimensional scaling are two-fold -to provide an information-rich display of the cocitation linkages and to identify the salient dimensions underlying their placement. ... Authors heavily cocited (because of their common subject or methodological interests) appear grouped in space. Authors with many links to others tend to be in central positions, while authors weakly linked, or with a few focused ties, will be placed in the periphery. In this way, “central” and “peripheral” research specializations, schools of thought, or other intellectual groupings can easily be seen, ... Dimensions are interpreted based on examination of the author and cluster placements.  ... The stress value reported for each solution (usually Kruskal’s Stress I or Stress II) and the proportion of variance explained (R Square in ALSCAL) are indicators of the overall “goodness of fit” of that point configuration.
c) Factor analytic techniques may be used to complement MDS and clustering displays. ... Essentially, they attempt to “explain” the interrelationships observed among the original variables through the creation of a much smaller number of “derived” variables or factors. In ACA, a factor is interpreted by the subset of authors loading on it - i.e., making substantial contributions to its construction. Essentially it reveals their underlying subject matter, as perceived by citers. ... ACA most commonly uses a principal components analysis, with an orthogonal (varimax) rotation of the extracted factors.
6) Interpretation and Validation: In ACA, interpretation and validation of results generally interact. Interpretation relies on discovering what the author clusters, factors, and map dimensions represent in terms of scholarly contributions, institutional or geographic ties, intellectual associations, and the like

van Eck, N. J., Waltman, L., Dekker, R. & van den Berg, J. (2010). A comparison of two techniques for bibliometric mapping: Multidimensional scaling and VOS. Journal of American Society for Information Science and Technology, 61(12), 2045-2061.

van Eck, N. J., Waltman, L., Dekker, R. & van den Berg, J. (2010). A comparison of two techniques for bibliometric mapping: Multidimensional scaling and VOS. Journal of American Society for Information Science and Technology, 61(12), 2045-2061.

本論文從學理與實驗兩方面比較MDS(Multidimesional Scaling)和VOS兩種資訊視覺化的維度縮減方法。作者認為從理論上來看,VOS是在計算Stress Function時以項目(items)間的相似程度(similarity)做為它們在圖形上對應點的接近程度(proximity)加權的一種特殊MDS方法,當兩個項目之間愈相似,它們映射在圖形上的點的接近程度便應該有愈大的加權。由於在實際的資訊視覺化應用上,大多數的項目之間沒有關聯,它們之間的相似程度經常被視為0,傳統的MDS方法也會運用這些項目之間的相似程度進行,導致映射的資料點形成一個接近於圓的圖形,觀測次數較多的項目較有可能會被映射到圓的中心,作者認為這樣的結果是失真的。作者以資訊科學(information science)的作者共被引關係等四種書目計量資料分別進行實驗,結果MDS的兩種相似程度的視覺化結果都相當接近圓,而VOS的結果比較令人滿意。
MDS has been widely used for constructing maps of authors (e.g., McCain, 1990; White & Griffith, 1981; White & McCain, 1998), documents (e.g., Griffith, Small, Stonehill, & Dey, 1974; Small & Garfield, 1985; Small, Sweeney, & Greenlee, 1985), journals (e.g., McCain, 1991), and keywords (e.g., Peters & Van Raan, 1993a, 1993b; Tijssen & Van Raan, 1989).
To determine similarities between items, co-occurrence frequencies are usually transformed using a similarity measure. Two types of similarity measures can be distinguished.
Direct similarity measures (Van Eck & Waltman, 2009; also known as local similarity measures, see Ahlgren, Jarneving, & Rousseau, 2003) determine the similarity between two items by applying a normalization to the co-occurrence frequency of the items. ... Various direct similarity measures are being used in the literature. Especially the cosine and the Jaccard index are very popular. ... We argued that the most appropriate measure for normalizing co-occurrence frequencies is the so-called association strength (e.g., Van Eck & Waltman, 2007b; Van Eck et al., 2006). This measure is also known as the proximity index (e.g., Peters & Van Raan, 1993a; Rip & Courtial, 1984) or as the probabilistic affinity index (e.g., Zitt, Bassecoulard, & Okubo, 2000).
Indirect similarity measures (also known as global similarity measures), on the other hand, determine the similarity between two items by comparing two vectors of co-occurrence frequencies. ... For a long time, the Pearson correlation has been the most popular indirect similarity measure in the literature (e.g., McCain, 1990, 1991; White & Griffith, 1981; White & McCain, 1998). Nowadays, however, it is well known that the Pearson correlation has some undesirable properties (Ahlgren et al., 2003; Van Eck & Waltman, 2008). A well-known indirect similarity measure that does not have these undesirable properties is the cosine.
The aim of MDS is to locate items in a low-dimensional space in such a way that the distance between any two items reflects the similarity or relatedness of the items as accurately as possible. The stronger the relation between two items, the smaller the distance between the items.
To determine the locations of items in a map, MDS minimizes a so-called stress function.  ... MDS determines the locations of items in a map by minimizing the (weighted) sum of the squared differences between on the one hand the transformed proximities of items and on the other hand the distances between items in the map.
Depending on the transformation function f, different types of MDS can be distinguished. The three most important types of MDS are ratio MDS, interval MDS, and ordinal MDS. Ratio and interval MDS are also referred to as metric MDS, while ordinal MDS is also referred to as non-metric MDS. Ratio MDS treats the proximities pij as measurements on a ratio scale. Likewise, interval and ordinal MDS treat the proximities pij as measurements on, respectively, an interval and an ordinal scale. In ratio MDS, f is a linear function without an intercept. In interval MDS, fcan be any linear function, and in ordinal MDS, f can be any monotone function.
The stress function in Equation 3 can be minimized using an iterative algorithm. Various different algorithms are available. A popular algorithm is the SMACOF algorithm (e.g., Borg & Groenen, 2005). This algorithm relies on a technique known
as iterative majorization.
The idea of VOS is to minimize a weighted sum of the squared distances between all pairs of items. The squared distance between a pair of items is weighed by the similarity between the items. To avoid trivial solutions in which all items have the same location, the constraint is imposed that the average distance between two items must be equal to one.
Under certain conditions, MDS and VOS are closely related. More specifically, the proposition indicates that VOS can be regarded as a kind of weighted MDS with proximities and weights chosen in a special way.
For each of the three data sets that we consider, three maps were constructed, one using the MDS-AS (direct similarity measure) approach, one using the MDS-COS (indirect similarity measure) approach, and one using the VOS (direct similarity measure) approach.
Experiment I: 405 authors publishing papers in 36 journals closely related to the Journal of the American Society for Information Science and Technology between 1999 and 2008.
Experiment II: 2079 journals that belong to at least one social science subject category.
Experiment III: 831 keywords that were automatically identified in the abstracts (and titles) of 7492 articles published in 15 operations research journals between 2001 and 2006.
A notable property of the maps produced by the two MDS approaches is that important items (i.e., items with a large number of co-occurrences) tend to be located toward the center of a map. This is especially clear in the case of the authors and keywords data sets. Many relatively unimportant items are scattered throughout the periphery of a map.
The VOS approach seems to produce maps in which important and less important items are distributed fairly evenly over the central and peripheral areas.
In various studies of the field of information science (e.g., Åström, 2007; White & McCain, 1998; Zhao & Strotmann, 2008a,b,c), it has been found that the field consists of two quite independent subfields. We adopt the terminology of Åström (2007) and refer to the subfields as information seeking and retrieval (ISR) and informetrics.
A distinction is sometimes made between ―hard‖ and ―soft‖ ISR research (e.g., Åström, 2007; Persson, 1994; White & McCain, 1998). Hard ISR research is system-oriented and is for example concerned with the development and the experimental evaluation of information retrieval algorithms. Soft ISR research, on the other hand, is user-oriented and studies for example users’ information needs and information behavior. The distinction between hard and soft ISR research is visible in all three maps.
Both (MDS) approaches have a tendency to locate the most prominent authors in the center of a map and less prominent authors in the periphery. Due to this tendency, the separation of subfields becomes more difficult to see.
In these maps(MDS-AS and MDS-COS), a number of prominent ISR authors (e.g., Spink, Wang, and Wilson) are located equally close or even closer to various informetrics authors than to some of their less prominent ISR colleagues. However, contrary to what the maps seem to suggest, there is in fact very little interaction between the prominent ISR authors and the informetrics authors. The relatively small distance between these two groups of authors therefore does not properly reflect the structure of the field of information science. The small distance is merely a technical artifact, caused by the tendency of the MDS-AS and MDS-COS approaches to locate important items in the center of a map.It follows from this observation that distances in maps constructed using the MDS approaches may not always give an accurate representation of the relatedness of items. Hence, in the case of the MDS approaches, the validity of the interpretation of a distance as an (inverse) measure of relatedness seems questionable.
The VOS map in ... does properly reflect the large separation between the prominent ISR authors and the informetrics authors. In this map, the interpretation of a distance as a measure of relatedness therefore seems valid.
This means that in the MDS-AS approach MDS is typically applied to similarity data that consists largely of zeros. MDS attempts to determine the locations of items in a map in such a way that for each pair of items with a similarity of zero the distance between the items is the same. In the case of similarity data that consists largely of zeros, it is not possible to construct a low-dimensional map with exactly the same distance between each pair of items with a similarity of zero. MDS can only try to approximate such a map as closely as possible. Our experiments indicate that the best possible approximation is a map with an almost perfectly circular structure.
From this point of view, one can say that the VOS approach distinguishes itself from the MDS-AS approach in that it does not give equal weight to all pairs of items. The VOS approach gives more weight to more similar pairs of items. It gives little weight to pairs of items with a low similarity. As mentioned above, similarity data is typically dominated by low values, in particular by zeros. ... In the case of the VOS approach, however, pairs of items with a low similarity receive little weight and therefore have little effect on a map. Because of this, the VOS approach does not produce circular maps.

Huang, Z., Chen, H., Guo, F., Xu, J. J., Wu, S., and Chen, W-H. (2004). Visualizing the expertise space. In Proceedings of the 37th Annual Hawaii International Conference on System Sciences (HICSS'04), IEEE Computer Society.

Huang, Z., Chen, H., Guo, F., Xu, J. J., Wu, S., and Chen, W-H. (2004). Visualizing the expertise space. In Proceedings of the 37th Annual Hawaii International Conference on System Sciences (HICSS'04), IEEE Computer Society.

information visualization/self-organizing map

本論文利用SOM及MDS等資訊檢索與文件處理技術將專家及他們的專長以視覺化的方式呈現在二維圖形上。這個研究的資料是台灣的597位商務與管理方面的學者,針對每位學者提供的研究領域(以國科會的分類,總共包括127個研究領域),在研究時分別對學者及研究領域建立特徵向量,用來產生專家地圖以及專長地圖。學者的特徵向量上每一個成分的二元值代表這位學者是否具有某項研究領域的專長,研究領域的特徵向量上每一個成分的二元值則是代表這項研究領域是否為某位學者的專長。最後將這些資料輸入SOM及MDS進行視覺化,進行MDS處理時兩個特徵向量間的相似程度是以Jaccard模式來進行估算。從結果的專家地圖上,可以發現具有相同與相近專長的學者被映射到相同或鄰近的節點上;在專長地圖上,有共同的理論或分析基礎或是共同的應用範疇的研究領域則會被映射到相同或鄰近節點上,形成群聚。
We focus on a basic form of expertise representation, in which experts are represented by a set of expertise fields. Due to the potential high dimensionality of such expertise data, we chose to examine two dimensionality reduction visualization techniques that have been widely applied in data and document visualization: the Self-organizing Map (SOM) and Multidimensional Scaling (MDS). We present two types of visualization results: the expert map and expertise field map, and provide initial analysis on the effectiveness of these visualizations to support expertise searching and browsing.
One type of set-level document visualization uses interactive scatter plots in different forms, which is also referred to as “dimensions and reference point systems” (Morse, Lewis and Olsen, 2000). Visualization techniques of this type attempt to display additional information about the retrieved documents and to group documents that share the similar characteristics. These characteristics may include the relationship between the documents and the query terms (Ahlberg and Shneiderman, 1994), predefined document attributes such as size, date, source and popularity (Hearst and Karadi, 1997;  Nowell, France, Hix, Heath and Fox, 1996), and user-specified attributes such as predefined topics (Olsen, Korfhage, Sochats, Spring and Williams, 1993).
A second category of techniques attempts to visualize inter-document similarities. This form of visualization is also referred to as “map systems” (Morse, Lewis and Olsen, 2000). There are four major techniques for inter-document similarity visualization: document networks (Thompson and Croft, 1989), physically based modeling techniques (Chalmers and Chitson, 1992), document clustering (Allen, Obry and Littman, 1993; Hearst and Pedersen, 1996) and geographic map metaphors (Chen, Schuffels and Orwig, 1996; Lin, Soergel and Marchionini, 1991).
Mockus and Herbsleb (2002) presented a system named “Expertise Browser” in the context of collaborative software engineering for change management systems. They embedded in some simple visualization elements such as the tree structure and other visual elements to present the expert attributes.
Our research explores this idea by focusing on a simple form of expertise database, where each expert is represented by a list of predefined expertise fields. Each expert can be represented as a binary vector, the elements of which correspond to the fields of expertise and the dimensionality is the number of predefined expertise fields. Each expertise field can also be represented as a binary vector, the elements of which correspond to the experts and the dimensionality is the number of experts in the data set. These representations adopt the vector space model of document representation and share the same high dimensionality characteristic. We chose two commonly used dimensionality reduction techniques for visualizing document space in the literature, the self-organizing map and multidimensional scaling, to generate map metaphors to visualize inter-expert and inter-expertise-field similarities.
The data set we used was resulted from an Internet survey on researchers in business and management fields in Taiwan. The survey was conducted by the National Science Council in Taiwan, and covered almost all the researchers in the business and management field in Taiwan.
The data set contained 597 researchers, who had selected their research interests or expertise from a two-level hierarchy of research fields. ... There were 127 second-level research fields and 2865 researcher-field combinations in the data set. ...  Each of the 597 researchers was represented by a binary vector with 127 elements, which corresponded to the research fields. The expertise similarity between two researchers was derived using vector similarity functions. We also had a dual representation for research fields, similar to the researcher representation. Each of the 127 research fields was represented by a binary vector with 597 elements, which corresponded to the researchers. In this case, similarities among research fields depended on the number of overlapping researchers. Such similarities may reflect the common theoretical/analytical foundations or closely related application domains of the research fields, based on the assumption that researchers typically work on closely related research fields.
The input to MDS is a square, symmetric matrix indicating relationships among a set of objects. Such matrices are usually either similarity or dissimilarity matrices. In the context of our research, a similarity matrix is formed based on the similarity scores of expert/expertise field pairs derived from the Jaccard’s similarity function (Jaccard, 1912).
We conducted a regression analysis to evaluate the general relationship between the researcher similarities and map distances. Researcher similarities were calculated using the Jaccard’s similarity function. A Euclidean distance function was used to calculate the map distances of researcher pairs. ... These statistics showed that SOM and MDS both preserved a large portion of the similarity information, although with certain degrees of distortion.
We observe from Figure 4 (expertise field map) that research fields having underlying similarities based on common theoretical/analytical foundations and application domains were grouped together. ... The expertise field map generated by our visualization techniques revealed meaningful grouping of research fields based on experts’ co-occurrence patterns in multiple research fields.

Zhu, D. and Porter, A. L. (2002). Automated extraction and visualization of information for technological intelligence and forecasting. Technological Forecasting & Social Change, 69, 495-506.

Zhu, D. and Porter, A. L. (2002). Automated extraction and visualization of information for technological intelligence and forecasting. Technological Forecasting & Social Change, 69, 495-506.

information visualization
本論文認為因為需要處理大量的文字資料、處理時要能快速以及呈現結果時需要生動並且能夠理解等三種需要,建議利用文字探勘(text mining)技術以及書目計量指標(bibliometric indicators)分析大量的文字資料庫,產生一系列的技術地圖(technology maps)和創新指標(innovation indicators)。其中產生各種技術地圖的技術包括在圖形上將資料項目對映到適合位置的MDS技術以及連結相關項目對映節點的路徑消除(path-erasing)演算法,並且本論文也建議使用詞語的共現資訊做為資料項目間相關程度的評估參考。
Three factors could enhance managerial utilization: capability to exploit huge volumes of available information, ways to do so very quickly, and informative representations that help manage emerging technologies.
Empirical analysis of emerging technologies poses a number of challenges to analysts. In particular, we note the need to:
1. digest enormous amounts of available information,
2. do so rapidly,
3. present findings vividly and understandably.
A third hard-earned lesson gained from our developmental experiences with ‘‘bibliometrics’’ (counting bibliographic activity) and ‘‘text mining’’ has been that TF-related results must be easily understood and must directly relate to a user’s perceived information needs.
This paper reports on efforts to address these three factors via partially automated processes to generate helpful knowledge from text quickly and graphically. We first illustrate a process to generate a family of technology maps that help convey emphases, players, and patterns in the development of a target technology. Second, we exemplify the generation of particular ‘‘innovation indicators’’ that measure particular facets of R&D activity to relate these to technological maturation, contextual influences, and market potential.
In sum, then, we seek to respond to these challenges—analyzing large text resources, rapidly, to generate compelling findings—to enhance TF (including competitive technological intelligence, technology foresight, etc.). Our approach, called technology opportunities analysis (TOA), seeks to facilitate this process by profiling search sets of bibliographic abstracts on technologies of interest.
The TOA process entails these main steps:
1. Search and retrieve text information, typically from large abstract databases.
2. Profile the resulting search set. VantagePoint applies a combination of machine learning, statistics, and natural language processing to yield what van Raan (1992) call a mix of ‘‘one-dimensional’’ descriptions (lists) and ‘‘two-dimensional’’ relationships (matrices). Profiling may focus on documents. Or, it may focus on concepts (e.g., principal components analysis (PCA) to group related terms as conceptual clusters). A third choice is a combination—seeking to link documents to concepts.
3. Extract latent relationships. VantagePoint applies iterative principal components analyses to uncover links among terms and underlying concepts.
4. Represent relationships graphically. Generation of ‘‘mapping’’ and ‘‘indicators’’ are elaborated in the following sections.
5. Interpret the prospects for successful technological development. This typically entails integrating the bibliographic search set analyses with expert domain knowledge (interviews).
We have developed a partly automated process to do so based on ‘‘co-occurrence’’ information. Co-occurrence is based on the pattern of terms occurring together in the records. If two terms occur together in the records more frequently than expected, there is a presumption of relationship between them. Terms can include authorship (also organizational affiliation, nationality) or ‘‘keywords’’ (subject index terms), or noun phrases generated from titles or abstracts using our natural language processing (NLP) routine (cf., Refs. [18,20]).
Effective visualization of the basic co-occurrence and correlation matrix information entails a sequence of analyses:
1) a new two-step multidimensional scaling (MDS) algorithm,
2) an improved path-erasing algorithm,
3) a routine to determine and display size (relative frequency of occurrence),
4) macros to create maps in VantagePoint, Microsoft Word or MS PowerPoint,
5) a routine to consolidate duplicate principal components (in the mapping process),
6) an algorithm to automatically name principal components,
7) an algorithm to cut off principal components to just include high-loading terms (the last three steps are needed for principal components maps; cf., Refs. [16,18]),
Our routine generates various maps, such as:
1. principal components map [represents the relationships among conceptual clusters];
2. keywords map [represents the relationships among frequently occurring subject index terms, title phrases, or whatever terms are chosen];
3. affiliations map [represents the relationships of affiliations’ research topics, based on terms they use in their documents—see Fig. 1];
4. authors map [analogous to affiliations map, but for individual researchers];
5. countries map [analogous to affiliations map];
6. sources (e.g., journals) map [analogous to affiliations map].
Fig. 1 shows an affiliations (organizations) map for the ‘‘Nanotechnology’’ topic. Displayed are the most prolific publishers abstracted in INSPEC for 1998. Along with the organizational name are shown the three keywords most frequently used in its publications in the search set. The size of a node reflects the number of publications. Positioning is determined using our MDS and path-erasing algorithm.
In essence, the challenge is to reduce n-dimensional (in this case, n equates to 40-dimensional since there are some 40 affiliations’ similarity being represented) to 2-D or 3-D. MDS is the generally favored approach to accomplish this. In MDS, an important parameter called stress is used to control its procedures. The process of generating a MDS map seeks the optimum location for each element in the map by minimizing the stress. ... We have devised a ‘‘step-by-step’’ search algorithm. This algorithm is effective at finding the global stress minimum, although it usually consumes more CPU time than the ‘‘steepest descent’’ algorithm.
Therefore, we have added an additional representational element, connecting links, based on a ‘‘path-erasing’’ algorithm. This is built on a proximity matrix among the elements. Its logic is as follows:
1. connect all elements in the proximity matrix together,
2. set a series of thresholds to erase the connecting lines one by one,
3. devise a suitable stop criterion.
The partially automated processes presented provide ‘‘value-added’’ knowledge from bibliographic text mining. The family of maps allows a user to gain an intuitive feel for R&D activity.
We suggest that development of routines to generate particular representations—technology maps and innovation indicators—automatically can enhance the applicability of text mining and bibliometrics to TF. ... However, scripting the production of these visualizations can facilitate provision of empirically based, vivid TF findings, in a timely manner, to inform decision making. That could dramatically increase the utilization of TF in management of technology

Moya Anegón, F., Herrero Solana, V. and Jiménez-Contreras, E. (2006). A connectionist and multivariate approach to science maps: the SOM, clustering and MDS applied to library and information science research. Journal of Information Science, 32(1), 63-77.

Moya Anegón, F., Herrero Solana, V. and Jiménez-Contreras, E. (2006). A connectionist and multivariate approach to science maps: the SOM, clustering and MDS applied to library and information science research. Journal of Information Science, 32(1), 63-77.

information visualization/self-organizing map
本研究以1992到1997年間LIS的17種期刊論文為研究資源,抽取論文和作者兩種單位的共被引關係,利用MDS和SOM兩種方法,並配合叢集分析(cluster analysis)產生科學地圖(science map)。作者認為這兩種方法都運用了維度縮減(dimensionality-reduction)的效用,並且具有互補的效果,研究結果發現不管是MDS或是SOM在作者共被引所產生的科學地圖上,大多都可以將圖形上映射的作者分為科學計量學(scientometrics)、引用關係研究(citationist)、書目計量學(bibliometrics)、傳播理論(communication theory)、資訊檢索的認知研究(cognitive information retrieval)和資訊檢索的演算法研究(algorithmic information retrieval),並且都屬於科學研究(science studies)的前三者在圖形上位置相當接近,兩種資訊檢索研究的距離也很相近。在論文共被引上,由於門檻較低,MDS和SOM的科學地圖上,除了原先的科學研究和兩種資訊檢索研究所組成的資訊科學研究以外,還包括圖書館研究與管理學兩個區域。
作者認為MDS產生的科學地圖可以保留論文(或作者)原先在高維度的距離關係,SOM則是保留了它們的型態(topology)關係。
The appearance of studies pertaining to library science reveals the relationship of this realm with information science. Especially significant is the presence of the management on the journal maps.
From a methodological standpoint, meanwhile, we would agree with those authors who consider MDS, the SOM and clustering as complementary methods
that provide representations of the same reality from different analytical points of view. ... This approach may be complemented with other kinds of representation based on network analysis.
Within the techniques of multivariate statistical analysis, three basic methods are included (Egghe and Rousseau, 1990):
1) cluster analysis,
2) principal component analysis (PCA), and
3) multidimensional scaling (MDS).
These methods are referred to as dimensionality-reduction methods because this function is to simplify what might at first appear to be a complex pattern of associations among many entities (Kinnucan, 1987).
SOM is based on the principle of the self-organization and grouping of n-dimensional vectors in a bidimensional space, and has been used to reduce dimensions in a wide variety of document spaces of diverse nature. ... According to Kaski (1997), the SOM presents four important properties for data exploration:
1) Ordered display. The characteristic help us to understand the underlying structures in the series of data.
2) Visualization of clusters. We are able to perceive the clustering density of the different regions of the map.
3) Missing data.
4) Outlier. This enables us to detect unusual cases caused by input errors or similar anomalies.
Journal co-citation mapping is potentially of interest both to the researcher studying the structure of scholarly specialities through the published literature and to the collection manager concerned with developing core journals lists, selecting journals and evaluating collections that serve particular research-oriented constituencies McCain(1991).
The main difference is that the SOM tries to present a locally corrected projection, whereas MDS attempts to preserve all the distances between the points. That is, MDS is distance preserving, while the SOM is topology-preserving.
Tijssen (1993) said that the mental representations of the experts, on an individual micro-level, and the bibliometric maps, on a macro-level, are inherently different.