顯示具有 SOM 標籤的文章。 顯示所有文章
顯示具有 SOM 標籤的文章。 顯示所有文章

2013年12月19日 星期四

Polanco, X., Francois, C., Lamirel, J. (2001). Using artificial neural networks for mapping of science and technology: A multi-self-organizing-maps approach. Scientometrics, 51(1), 267-292.

Polanco, X.,  Francois, C.,  Lamirel, J. (2001). Using artificial neural networks for mapping of science and technology: A multi-self-organizing-maps approach. Scientometrics, 51(1), 267-292.

information visualization/self-organizing map

本篇論文提出一個利用multi-SOM為基礎的文件主題知識介面,以一個SOM表現一種描述文件方式的詞語資料(例如:論文中作者姓名或是研究方法相關的詞語),並且利用SOM之間的對映關係,將multi-SOM串聯起來提供研究者在分析時的參考。這篇論文討論了(1)將文件在多維度上的距離關係盡量保留映射到二維平面上的關係,(2)出現在較多文件的詞語會在圖形上佔有較多個節點(也就是較大的面積)等SOM的特性,並且建議使用SOM作為分析參考時的程序,包括(1)在SOM產生後檢視主題的合理性,(2)解釋SOM圖形所表現的訊息,和(3)將不同的SOM圖形串聯起來。這篇論文的研究以植物相關的1843筆專利文件為探討對象,並且以724個詞語作為索引。

According to Kohonen (1997, p. 86), one might say that the SOM is a non-linear projection of the probability density function p(x) of the high-dimensional input data vector x onto the two-dimensional display (i.e., a map). ... The self-organizing map (SOM) gives central attention to spatial order in the clustering of data. The purpose is to compress information by forming reduced representations of the most relevant features, without loss of information about their interrelationships.
In the quantitative studies of science, the Kohonen self-organizing maps have been used for mapping scientific journal networks (Campanario, 1995), and also author cocitation data (White et al., 1998).
In the SOM, the competitive learning means also that a number of nodes is comparing the same input data with their internal parameters, and the node with the best match (say, “winner”) is then tuning itself to that input, in addition the best matching node activates its topographical neighbours in the network to take part in tuning to the same input. More a node is distant from the winning node the learning is weaker.
Like any unsupervised clustering method, the SOM can be used to find clusters in the input data, and to identify an unknown data vector with one of the clusters. Moreover, the SOM represents the results of its clustering process in an ordered two-dimensional space (R2). A mapping from a high-dimensional data space Rnonto a two dimensional lattice of nodes is thus defined. Such a mapping can effectively be used to visualise metric ordering relations of input data.
The SOM takes a set of documents in our case patents as input data (x), each patent is represented by an N-keywords vector (x∈Rn), and maps them onto nodes of a two-dimensional grid (n∈R2).
The main properties of such self-organizing maps are the following: “First, the distance relationships between the input data are preserves by their images in the map as faithfully as possible. While some distortion is unavoidable, the mapping preserves the most important neighbourhood relationships between the data items, i.e., the topology of their distribution. Second, the map allocates different numbers of nodes to inputs based on their occurrence frequencies. If different input vector appear with different frequencies, the more frequent one will be mapped to larger domains at the expense of the less frequent ones” (Ritter and Kohonen, 1989, p. 246).
The algorithm is based on three computational levels: the winning node selection, the unsupervised learning and neighbouring definition, and the control mechanisms of the unsupervised self-organizing algorithm.
Due to the fact that there is obviously no absolute strategy for achieving that goal, the choice has been to implement two different kinds of strategies that could be indifferently used during the map interactive consultation phase. They are respectively called the clusters vector driven strategy and the document vector driven strategy. ... The clusters vector driven strategy consists of attributing to each cluster a name that represents the combination of the labels of the components having the maximum values in its vector. This strategy is well-suited in highlighting for the user the main themes described by the map. ... The document vectors driven strategy consists of attributing to each cluster a name that represents the combination of the labels of the components having the maximum values in either the vector of the most representative member of the cluster or the average document vector computed from all the cluster member vectors.
In order to reach that goal the task that the system operates is to reduce the number of cluster (i.e., the number of nodes) of the map in a coherent way. The method consists in starting from the original map and introducing new clustering levels of synthesis (i.e., maps) by progressively reducing the number of nodes. Since the original map has been build on the basis of a 2D square neighbourhood between nodes, the transition from one level to another is achieved by choosing a new node set in which each new node will represent the average composition of a square of four direct neighbours on the original level. ... This procedure has the advantage of preserving the original neighbourhood structure on the new generated levels. Moreover it ensures the conservation of topographic properties of the map nodes vectors, and consequently the conservation of the closeness of the nodes areas in the generalized maps.
Our empirical example is a set of 1843 patents about vegetal transgenic technology indexed by 724 keywords, and recorded in the period 1978–1997.
In comparison with the standard mapping methods, as such principal component analysis or multidimensional scaling, the advantage of the multi-map displays is the inter-map communication mechanism that Multi-SOM environment provides to user. Each map is representing a viewpoint. Each viewpoint is representing a subject category. The inter-map communication mechanism assisted the user to cross information between the different viewpoints.
On the computer system side, the following task was to build the maps representing the different viewpoints, using the map algorithm. The second was to use the inter-map communication for achieving thematic querying.
The system provides to human analysts a first level analysis with its unsupervised learning approach for extracting from the data the features that the maps display. The second level of analysis is constituted by the tasks that the human analysts should achieve. These tasks can be organized in three successive stages.
First stage: Validation of the Plant Map. Verifying if the clusters of an area really represent the plant instead of the usual host of a pathogen. Examining the cluster positions on the map; Observing the cluster vectors; Considering the relative size of the associated areas.
Second stage: Interpretation of the results. This asks strongly the background knowledge of the domain experts. ... The domain expert also detected ambiguous clusters on the map.
Third stage: Thematic queries using inter-map communication process. The activation of the whole area associated to a plant or a plant group. The use of the inter-map communication of the activity applying a possibilistic parameter along with a bias from the activated documents to the clusters of the target maps. The analysis of all the activated clusters on each of the target maps (i.e., target viewpoint).
Two goals could be achieved at least by the graphic user interface. The first is the working interface for elaborated clean and final results. The second is the visualisation of the clean and final results with browsing and querying functionalities.
The model that this multi-map environment provides is certainly the map but in its original extended version of intercommunication between multiples maps. Each map represents a particular viewpoint extracted from the data. These viewpoints are related either by the problem to be solved, or by the intercommunication mechanism between the maps. We have exposed both the map generation and their intercommunication mechanism. We finally showed how this clustering and mapping environment gives assistance to users in some watching intention.
A reason to use ANNs in quantitative studies of science and technology is their capability to create “higher abstractions from raw data completely automatically. Intelligence in neural networks ensues from abstractions, not from heuristic rules or manual logic programming” (says Kohonen, 1997, p. 65).
The maps play the role of strategic indicators because they provide a comparison way for evaluating the relative position of themes onto an ordered space.

Campanario, J. M. (1995). Using neural networks to study networks of scientific journals. Scientometrics, 33(1), 23-40.

Campanario, J. M. (1995). Using neural networks to study networks of scientific journals. Scientometrics, 33(1), 23-40.
information visualization/self-organizing map
本研究以自組織映射圖呈現期刊相互引用的網絡結構,作者認為由於Kohonen自組織映射圖本身的數學形式將使得愈緊密連結的期刊映射在愈靠近的位置,並且自我引用較多的期刊在圖形上佔有較大的面積。作者進行了四個資料集的期刊引用網絡研究,包括19種化學物理期刊和20種傳播學期刊,其餘兩個資料集是不同年份的社會學期刊,分別包括11種及13種期刊,用來比較領域在不同時間區間內是否產生變化。每一種期刊表示成一個n維的特徵向量,n是資料集內的期刊種類,特徵向量上的每一個成分對應到該期刊引用另一個期刊的次數。
If one paper cites an earlier publication, they bear a conceptual relationship. The references given in a publication link that publication to previous knowledge. In the network context, information is reconceptualized in terms of social linkages and shared meanings. According to Small and Garfield(1985), citation indexes, showing millions of interconnections annually among hundreds of thousands of scientific articles and books, seem ideally suited for deriving natural maps of the scientific landscape.
The results of studies carried out with the above methodologies have been used to identify science maps (Small, Sweeney, and Greenle, 1985; Small, 1993), maps of disciplines (Garfield, 1986), research fronts (Dixon, 1989), scientific journal networks (Midorikawa, 1983; Pinski, 1977; Saito, 1990), epistemic and conceptual networks(Leydesdorff, 1991; van Raan and Tijssen, 1993), invisible colleges (Lievrouw, Rogers, Lowe, and Nadel, 1987) or author networks (McCain, 1986) and to establish the rank of journals in a given network (Doreian, 1987; Hummon and Doreian, 1989; Doreian, 1994; Bonitz, 1990).
Journals are a central institution of science because they are the primary formal channels for communicating theories, methods and empirical results to the readers of those journals (Rice, 1990).
Four sets of journal-to-journal citation data were used to apply Kohonen's map algorithm to the study of networks of scientific journals. ... Data are for citations among 19 chemical physics journals pooled for 1981 (data set I), 20 communication journals pooled from 1977 to 1985 (data set II), 11 sociology journals pooled from 1970 to 1970 (data set III) and for citations among 13 sociology journals from 1975 to 1976 (data set IV). Data sets III and
IV include data referring to the same journals (for different years) with some new journals added in data set IV. This makes it possible to compare the results of two different years that are made up of almost the same journals.
To perform the computations, each journal was coded as an n-component vector (n representing the number of journals in a given set). Each component of the vector is the number of citations given by each journal to each other journal. The input vectors were normalized to allow the algorithm to normalize the weights.
The figures show the distribution of relational space among the journals. Because of the mathematical formalism of the Kohonen maps, the most closely linked journals are located close to each other. In addition, domains occupied by the journals with a large number of self-citations tend to be greater.
Most of the multidimensional statistical methods use a symmetrical matrix of relations among cases to define a distance. However, the cross-citation matrix among journals is asymmetric, as noted earlier. This fact reflects the hierarchical structure of journal relationships: some journals are subordinate to others (Leydesdorff, 1986). This problem is sometimes overcome by computing some kind of correlation in order to obtain a symmetrical cross-citation matrix. However, this transformation causes the loss of the hierarchic quality of journal interrelations. This is manifested in the domain map in which, sometimes, a given journal activates more cells than other closely linked journals.

Guerrero Bote, V. P., de Moya Anegón, F. and Herrero Solana, V. (2002). Document organization using Kohonen's algorithm. Information Processing and Management, 38, 79-89.

Guerrero Bote, V. P., de Moya Anegón, F. and Herrero Solana, V. (2002). Document organization using Kohonen's algorithm. Information Processing and Management, 38, 79-89.

information visualization/self-organizing map

本研究以Self Organizing Map方法,將LISA資料庫中八類描述語(Acquisitions, Artificial Intelligence, Business Management, Computerized Information Storage and Retrieval, Conferences, Periodicals, WWW)的202筆摘要進行組織。結果具有鄰近的節點大多是具有相同描述語的摘要,而且鄰近區域也可以找出關連,例如Computerized Information Storage and Retrieval在產生的兩個圖形上所佔的區域與Artificial Intelligence和WWW的區域都相鄰。

The Kohonen's model is capable of performing a topological organization of the inputs presented to it.
This type of network has recently been used in documentation for the analysis of domains (White, Lin, & McCain, 1998), for textual data mining (Lagus, Honkela, Kaski, & Kohonen, 1999), to extract semantic relationships between words from their contexts (Honkela, Pulkki, & Kohonen, 1995, Ritter & Kohonen, 1989), and in particular to generate topological maps of sets of documents, even labeling the zones of influence of each word or term (Kohonen et al., 1999a; Kaski, 1999, Lagus & Kaski, 1999, Moya, Herrero, & Guerrero, 1998; Moya Anegón et al., 1999, Chen, Houston, Sewell, & Schatz, 1998; Lin, 1997; Huerrero Bote, 1997; Orwig, Chen, & Nunamaker, 1997; Lin, Soergei, & Marchionini, 1991).
In the learning process, as well as clustering the inputs, the Kohonen network generates a topological organization of those clusters. When we apply this to documentation the result will be the creation and organization of clusters in a manner that those which are topically close will also be close in the network. We may use this to expand the query, or rather the results: once one has found the cluster that best fits the query, one may extend the activation to those which are topologically close.
In some of these case, as well as performing a document classification, one determines for each node which unitary term vector produces the greatest activation. One may thereby generate each term's zone of influence, providing a graphical view of the database on which one could even select the zone that one wants to visit.

Linton, J. D., Himel, M. and Embrechts M. J. (2009). Mapping the structure of research: Business and Management as an exemplar. Serials Review, 35, 218-227.

Linton, J. D., Himel, M. and Embrechts M. J. (2009). Mapping the structure of research: Business and Management as an exemplar. Serials Review, 35, 218-227.

information visualization/self-organizing map

本論文以202種商學與管理類的期刊為研究對象。每一種期刊蒐集約200筆摘要,統計詞與詞對在所有摘要內的出現次數,選取出現次數較多的300個詞與163個詞對做為摘要的特徵,建立202種期刊的特徵向量,訓練自組織映射圖,並且進行映射。在產生的自組織映射圖上,領域相關的期刊可以被映射到相同或相鄰的節點上,如果期刊之間的映射距離較遠,它們之間的關係較小。作者認為這種方法所建立的自組織映射圖可以提供新進研究人員在選擇期刊上的參考,做為研究人員升等的評估參考,了解跨學科的研究領域的特性,並且提供大量資訊較佳的組織方法。
The relationship of different journals to each other is of interest to management academics and librarians for a number of reasons:
1) New researchers (students and junior faculty) often find it difficult to determine which journals offer a good fit with their interests because they are new to the field.
2) Evaluation of promotion and tenure is often complicated by lack of common agreement and sufficient domain knowledge for assessing the relevance and quality of the journals a candidate has published in.
3) Better understanding of the interdisciplinary nature of a field and the relative distance and proximity of different subfields is essential.
4) Demonstration of techniques and methodology to better organize large amounts of information without having to find individuals with suitable domain knowledge is useful and does not risk selection biases that can result from reliance on individuals.
5) Selection/cancellation of journals in a collection depends on many metrics. ... These techniques can assist in identifying the fit of journals that are candidates for the addition or cancellation process in a serials collection.
Parameswaran and Sebastian (2006) indicate that the problems with objective ranking studies include the following:
1) bigger journals may have higher citation scores because there is greater potential for citation;
2) journals connected with professional associations have a large default subscriber base;
3) all references are not equally important;
4) authors tend to cite more from their own culture;
5) citations are countedwhether they involve praise or criticism;
6) citation analysis does not capture influence that extends beyond academia;
7) some seminal works are so well known that authors no longer feel the need to reference them.
Method used in this study
202 business and management journals selected from the “Business,” “Business-Finance,” and “Management” categories of the Social Sciences Citation Index, the “Operations Research and Management Science” category of the Science Citation Index and the "Top 40" journal list selected by the Financial Times.
The 200 most recent abstracts were downloaded for each selected journal.
The 300 most frequently occurring words and 163 most frequently occurring word couplets, with the condition that the words that lacked specificity to business and management research were eliminated, were selected as a “dictionary” of common terms for the business and management field. Each journal is described by a vector offering the relative occurrence of the commonly occurring technical words and word pairs. Once this process was completed for all of the journals in the study , the set of vectors describing the journals was entered into a Kohonen self-organizing map (SOM) program.
 Results of the Mapping
The journals are all arranged based on the relative use of commonly used technical words. Journals that sit in the same cell are very similar to each other. Journals that are sitting in adjacent cells have clear similarities in relative frequency of common technicalwords.While there are some similarities amongst
journals that have a cell situated in between them, once journals are further apart than this, the degree of similarity declines rapidly.
Another area in which interesting insights are offered is into the structure of interdisciplinary research. In some cases, journals associated with different traditional disciplines sit in adjacent cells. This relationship suggests that the journals may offer a bridge between these fields.
The self-organizing map has substantial utility for management researchers and practitioners in the following ways:
1) This tool offers new researchers a way of accelerating the build up of domain knowledge regarding which journals are and are not a fit with their research interests and agenda.
2) The SOM can assist in bridging the lack of agreement and insufficient domain knowledge that adds great variability into the assessment of the relevance and quality of journals into the tenure and promotion process.
3) This approach can create a better understanding of the interdisciplinary nature of the field and the relative distance and proximity of different fields. The SOM can also be used to identify journals that have a role in acting as a bridge spanning two or more disciplines or subfields.
4) The process demonstrates techniques and methodology to better organize large amounts of information without having to find individuals with suitable domain knowledge and thus risk selection biases that can result with relying on individuals.

Lin, X. (1997). Map displays for information retrieval. Journal of the American Society for Information Science, 48(1), 40-54.

Lin, X. (1997). Map displays for information retrieval. Journal of the American Society for Information Science, 48(1), 40-54.

information visualization/self-organizing map

本研究分析瀏覽做為資訊檢索方式的適用性以及各種以瀏覽為基礎的視覺化組織格式。

資訊檢索包含搜尋及瀏覽兩種方式。若是要將瀏覽應用做為一種資訊檢索的方式,必須考慮 1)資訊項目需要具有良好的組織結構,2)使用者有需要探索他們不熟悉的集合內的資訊項目,3)使用者不了解集合內的資訊組織並且希望有較低認知負荷的探索方式,4)使用者對表達他們的資訊需求有困難和5)使用者能夠識別他們想要的資訊,但很難描述它們。

為了在資訊檢索服務提供瀏覽方式,需要將大量的資訊項目進行視覺化組織,並且希望能夠保留資訊的結構與關係,以提供使用者能有效地運用他們的視覺能力。因此本研究也比較階層式、網絡式、散佈式和地圖式等四種視覺化組織形式的特性與優缺點。
1) 階層式顯示的優點在於能夠藉由分層、分支與分群等方式簡化複雜的資料結構,可以同時表現出資訊的全體性質與區域性質,而且可以將觀看者的注意力集中在適當的廣度上。但若是資訊本身的結構不是階層式時,階層式結構往往過度簡化;此外,若是資訊空間較大時,較難產生與顯示階層式結構;並且使用者在選擇接下來要瀏覽的分支時需要較大的認知負荷。
2) 網絡式顯示以圖形上的節點代表資料項目,使用者在瀏覽時可以循著節點之間的連結線找到相關的資料項目。相較於階層式顯示,網絡式顯示能夠表示更複雜的結構,應用的範圍較廣。但也因為如此,當網絡結構較為複雜時,使用者可能不容易理解。
3) 散佈式顯示運用映射演算法將高維度的資料對應到二維的圖形平面上,並且這樣的映射過程必須使資料間原先的距離關係得以盡量保存在對應後的結果上。散佈式顯示非常適用於表現資料的整體型態,並且能夠展現出資料的意義結構。
4) 地圖式顯示劃出成若干區塊,每個區塊代表一個可能的主題(由一組相關的詞語組成),區塊的尺寸大小表現主題的重要性(愈出現的詞語在地圖上佔有愈大的區塊,而其出現的),區塊之間的距離遠近則表現出主題之間的關連程度,距離愈近的區塊表示主題之間愈相關。作者並認為地圖式顯示兼具前三者顯示方式的優點,也就是可以包含階層式顯示的階層的叢聚,網絡式顯示的相關連結,以及散佈式的空間映射方式。有別於前三者,地圖式顯示方式並不直接輸出最後文件資料映射到圖形上的結果,而是產生一個關聯式網路(associative network)。

本研究以三個文件資料集做為範例,以SOM做為地圖式顯示的方法,並且運用不同的索引(文件特徵)方式代表這些資料集內的文件資料。
(1) 311篇多語言資料檢索(multilingual information retrieval)主題相關的論文,以論文題名中出現的85個詞語為特徵,特徵向量上每個成分為是否有出現的二元值。結果發現地圖上的區域能代表論文的主題,區域的大小與主題出現的論文數量有關,較常出現的主題佔有較大的面積,相連的區域表示這些主題曾經共同在論文內出現。
(2) 660篇研究者個人蒐集的論文,以論文題名、關鍵詞及摘要出現的1472個詞語為特徵,特徵向量上每個成分為詞語在該篇論文資料的出現次數乘上詞語的倒文件頻率(inverse document frequency)。結果發現地圖上的區域能表示研究人員感興趣的主題,區域的面積愈大表示研究人員對這個主題愈重視;當論文映射到自組織映射圖上相對應的位置時,也可以發現論文的分布情形。
(3)143篇1990-1993年間SIGIR的會議論文,以論文題名出現的154個詞語為特徵,特徵向量上每個成分為詞語在該篇論文的題名、關鍵詞及摘要等資料的出現次數。產生的地圖作為資訊檢索的互動介面,可以增減自組織映射圖上出現的詞語數量,並可以點選圖上的任一區域檢索相關的論文。
Computers are expected to be used to reveal associations and properties of electronic information to allow people to use their visual capabilities for information seeking (Veith, 1988) .
The map display attempts to show both contents and semantic structures of a document space by mapping major concepts and documents of a document space to a two-dimensional map. It preserves, as faithfully as possible, document semantic relationships and reveals these relationships through various visual components of the display. 
Most users have difficulty specifying their needs by a specific query formulation; even if users are successful in doing so, systems have difficulty retrieving all relevant documents without overwhelming the users with irrelevant documents. The issue of precision/ recall has been a bottleneck for retrieval systems: Retrieving more relevant documents (high recall ) is often at the price of getting more irrelevant documents (low precision) .
Visual displays that show terms and document relationships and reveal underlying structures of the document space will be such browsing aids that will relax demands on the performance of retrieval mechanisms and query generations. Such displays will allow the user to interact and browse a large quantity of search results in a limited display space.
Browsing is a direct application of human perception for information seeking, both in the electronic and non-electronic environment (Chang & Rice, 1993) . Browsing is explorative; it is an interactive process in which one will scan large amounts of information, perceive or discover information structures or relationships, and select information items through focusing one’s visual attention.
In relation to information retrieval, browsing is particularly useful when:
1) There is a good organizational structure and related information items are often located near each other (Thompson & Croft, 1989).
2) Users are not familiar with the content of the collection and they need to explore the collection (Motro, 1986).
3) Users have less understanding of how information is organized in the system and they prefer to take a low  cognitive load approach to explore the system (Marchionini, 1987) ,
4) Users have difficulties in articulating their information needs (Belkin, Oddy, & Brooks, 1982).
5) users look for information that is easier to recognize than to describe (Bates, 1986) .
Some techniques that researchers have explored to support browsing for information retrieval include:
1) displaying a dynamic hierarchical information structure (Frei & Jauslin, 1983) ,
2) providing an overview map of the information space (Halasz, Moran, & Trigg, 1987) ,
3) providing a neighborhood map for each item (Thompson & Croft, 1989) ,
4) showing both a miniature of the entire information space and a detailed local map (Beard and Walker, 1990) ,
5) distorting the display so that the center of focus will be shown in more detail than other areas—the fish-eye views (Furnas, 1986) , and
6) supporting interactive functions such as zoom in, zoom out functions so that the user can select different level of details to display (Schatz & Caplinger, 1989) .
A central issue of organizing information for visualization is what formats and features of visual displays will help to organize large amounts of information to reveal information structures and to support effective use of human visual capabilities.
Hierarchical displays simplify complex data structures and separate data into different levels, branches, or clusters. These functions help to represent both global and local views of data, to utilize the display screen effectively, and to direct the viewer’s attention to the appropriate level of generality.
Cutting, Karger, Pedersen, & Tukey, (1992) showed that hierarchical clustering could be an effective information access tool, particularly for browsing.
These disadvantages of hierarchical displays include (1) oversimplification of structures for certain data, particularly for those that are more appropriate to be represented by structures other than a hierarchy, (2) difficulty in generating and displaying hierarchical displays for large information spaces, and (3) increased cognitive load for users who are forced to make selections among the hierarchical branches, especially when the whole hierarchy is not displayed on the screen.
Network displays show associative structures on the screen and let the viewer follow the links to browse items represented by the nodes. ... They can represent more general and complicated structures than hierarchical displays can. ... However, if all the relationships in a complex document space are displayed in a network, the network display simply becomes a network maze. The network displays thus often present more information than the user can immediately comprehend (Beard & Walker, 1990) .
Scatter displays refer to the graphical (dotted) image resulting from mapping high-dimensional data to a two-dimensional visual space. ... Most of the scatter displays are generated automatically by mapping algorithms. Because the mapping is usually driven by an error-minimum process or by the principle of finding a display configuration whose overall layout most closely matches the structure of the given data, the mapping creates a spatial orientation that reflects the overall layout of underlying data.
Scatter displays are very useful in revealing underlying data structures of statistical data (Tufte, 1983) . In particular, scatter displays can also be used to reveal semantic or intellectual structures embedded in statistical data.
Among the three display formats reviewed, scatter displays most faithfully reflect underlying data structures. In a scatter display, the viewer is not constrained to follow predetermined links as in the network display or to follow a rigid hierarchical structure in the hierarchical display. However, this lack of regularity in the scatter display also poses problems for the viewer trying to discover the underlying structure. In this respect, the scatter display particularly needs the help of other context or interactive probes such as verbal labeling or mouse sensitive areas.
Compared to the physical space, the document space is much less clearly defined in terms of its measurement, its dimensionality, and its semantic relationships, all of which largely depend on the selected indexing process. ... It would be difficult to have a map that is a ‘‘true’’ representation of the document space like the geographical map is for the physical space. ...  The map displays should also provide rich visual information, and be able to present dynamic displays at different detail levels to allow the user to interact with the underlying information.
The map display was designed to provide the advantages of mapping, linking, and clustering as in the scatter displays, network displays, and hierarchical displays reviewed earlier.
The mapping algorithm selected will keep the display structure as similar to the underlying data structure as possible.
With an appropriate indexing, Kohonen’s feature map algorithm can be used to ‘‘survey’’ contents of a document space, to ‘‘detect’’ semantic relationships of terms and documents, and to generate map displays that will show both contents and semantic relationships of documents.
Kohonen’s feature map algorithm takes a set of input objects, each represented by an N-dimensional vector, and maps them onto nodes of a two-dimensional grid.
The mapping procedure is a recursive learning process of the following:
1) Select an input vector randomly from the set of all input vectors,
2) find the node (which is also represented by an N-dimensional vector called weights) closest to the input vector in the N-dimensional space,
3) adjust weights of the node (called the winning node) , so that it will more likely be selected again if this input is presented later,
4) adjust the weights of those nodes within a neighborhood of the winning node, so that nodes within this neighborhood will have similar weight patterns.
This process goes through many iterations until it converges, i.e., the adjustments all approach zero.
To ensure its convergence, two control mechanisms are imposed.
The first is the updating parameter. It approaches to zero as the number of iterations increases.
The second is the neighborhood structure that shrinks gradually during the process. A large neighborhood will achieve ordering and a small neighborhood will help to achieve a stable convergence of the map (Kohonen, 1989) . By beginning with a large neighborhood and then gradually reducing it to a very small neighborhood, the feature map achieves both ordering and convergence properties.
Early applications of the algorithm mostly demonstrated that the feature map could preserve metric relationships and the topology of input patterns.
I. A Map Display for a Retrieved Set of Documents
This example used a set of documents retrieved by a search done on INSPEC database in DIALOG for the topic of multilingual information retrieval. The set contains 311 documents. The indexing for this document set was based on titles only. ... As the result, 85 terms were retained to index the document set. A vector of 85 dimensions was created for each document, where a component was a ‘‘1’’ if the corresponding term occurred in the document title and a ‘‘0’’ otherwise. The document vectors were used as input to train a feature map of 85 input features and 10 by 14 output nodes arranged in a grid.
Results:
The areas on the resulting map can be seen as concept areas(more precisely, word areas) . 
The size of the areas corresponds to the word occurrence frequencies.
The neighboring relationships of areas indicate frequencies of the co-occurrence of words represented by the areas.
A Map Display for a Personal Collection
The second example is a map display for a personal document collection. The collection contained 660 documents, which were accumulated over many years as a by-product of a researcher’s research activities. ... The indexing for this collection was fulltext-based—every word in the titles, keywords, and abstracts was used. After the stopword-removing and stemming procedures, and the elimination of the most-frequently and the least-frequently occurring terms, 1,472 terms remained in the indexing list. To create the indexing vectors, weights of each term were computed based on both the term frequency and the inverse document frequency. The 660 document vectors of 1,472 dimensions were then used as input to train a 10 by 14 Kohonen’s feature map of 1,472 input features.
Results:
The map display generated shows the researcher’s major research areas and the relationships of these areas.
The size of areas, corresponding to the frequencies of the words, indicates relative importance of the areas to the researcher (the more often a word appears in the personal collection, the more likely the word will correspond to a large area in the space) .
The neighboring relationships, corresponding to the frequencies of co-occurring words, reflect degrees of word associations as derived from the researcher’s collection.
When each document in the collection was mapped to a position on the display, the document distribution over the map display can also be shown.
The map can also reveal migration of the researcher’s interest over time.
A Map Display for Conference Proceedings
The third example is about documents from 1990–1993 SIGIR conference proceedings. These proceedings contain 143 documents. The indexing terms for this collection were collected from titles only, but the weights of terms were computed based on the term frequency in titles, keywords, and abstracts. ... After the same stopword-removing, stemming procedures, and elimination of the most- and the least-frequently occurring terms, 154 terms were used to index the collection, resulting in 143 vectors of 154 dimensions. These vectors were then used to train a 14 by 14 feature map of 154 input features.
As a mapping tool, the feature map has the properties of economic representation of underlying data and their interrelationships ( in these examples, the feature map self-organizes major terms selected from hundreds or thousands of indexing terms to represent the document spaces) .
As a visualization tool, the feature map produces rich geographical features that can be used for visual inferences. The algorithm generates an associative network as the output, rather than the direct mapping of the input. This makes it easy to implement various interactive tools and provide different ‘‘views’’ of the underlying information.
Finally, that the algorithm allows classification of any input to more than one location is certainly beneficial to information retrieval.
A major challenge to the success of document mapping is how to evaluate map displays and how to compare different map displays. ... This result suggests that comparisons of map displays need to be done on how the map displays help the user locate documents, not just how they look. It is quite possible to have different organizations of map displays that can provide the same level of access to a document space.

Smith, K. A. and Ng, A. (2003). Web page clustering using a self-organizing map of user navigation patterns. Decision Support Systems, 35, 245-256.

Smith, K. A. and Ng, A. (2003). Web page clustering using a self-organizing map of user navigation patterns. Decision Support Systems, 35, 245-256.

information visualization/self-organizing map

本研究利用自組織映射圖技術,將235個 Monash University, School of Business Systems的網頁進行叢集。本研究利用網頁伺服器上儲存的記錄檔做為自組織映射圖在訓練及映射時的參考資料,首先將記錄檔依據記錄上的使用者和時間資料,劃分成8054個Transactions,再以K-means叢集方法,根據每個Transaction裡的網頁,將Transactions歸類成9類,統計235個網頁在9個Transaction分類上的數目,建立代表網頁的特徵向量。由於在結果的自組織映射圖上,相同目錄的網頁會被映射到相同或鄰近的節點上,由此可見,利用自組織映射圖技術以及使用者存取網頁的資料可以用來對相關的網頁進行歸類與視覺化。
For the system to accurately reflect the needs of users, the organization of the web documents should also take into account the feedback from users. While it is useful to have a system to organize the web pages in a content-driven manner, it may be more advantageous to organize the web pages in a web-user oriented manner. After all, the web documents are organized so that humans can search in a more effective and efficient manner.
The authors have developed the prototype of the LOGSOM system based on the access logs for September of 1999 from the Monash University, School of Business Systems web server. There are 170,515 entries in the web log indicating the date, time, and address of the requested web pages, as well as the IP address of the user’s machine. ... The original server logs are formatted, cleansed, and then grouped into meaningful transactions before being mapped onto the self-organizing map.
Following Cooley et al. (1999), the authors group the data into meaningful transactions. The authors define a transaction as a
set of web pages requested by a user in a particular session. ... For the examined web log, the number of transactions m= 8054 and the number of URLs n = 235.
The number of inputs of the SOM will need to be equivalent to the number of transactions, and because this number is so large, it will not be feasible with this data. ... By using the K-means clustering algorithm, we cluster the transactions into nine groups. The number K=9 is chosen arbitrarily. ... Thus, after the dimension reduction, it consists of 235 URLs X 9 transaction-groups.
The distance between nodes on the resulting map indicates the similarity of the web pages, measured according to the user navigation patterns. LOGSOM provides a visual tool to enable users to see the relationship between web pages based on the usage patterns of web users similar to themselves. LOGSOM also provides an analysis tool for web masters and web authors to better understand the interests of visitors to their pages, and identify potential referring pages.

White, H. D., Lin, X., Buzydlowski, J. W. and Chen, C. (2004). User-controlled mapping of significant literatures. Proceedings of the National Academy of Science of the United States of America, 101, 5297-5302.

White, H. D., Lin, X., Buzydlowski, J. W. and Chen, C. (2004). User-controlled mapping of significant literatures. Proceedings of the National Academy of Science of the United States of America, 101, 5297-5302.

本研究使用尋徑者網路(pathfinder networks, PFNET)和自組織映射圖(self-organizing maps, SOM)兩種維度縮減(dimension reduction)技術做為PNAS期刊論文檢所的圖形化介面,輸入一個詞語或一位作者,產生這個查詢與其相關的24個詞語或24位作者的圖形。以Gene Frequency與這個主題的重要作者Slatkin做為查詢所產生的PFNET與SOM,提供Slatkin檢視,他認為產生的圖形很容易解釋。本研究並說明與比較了PFNET和SOM作為資訊視覺化界面的特點。

information visualization

Our data are the contents of PNAS for 1971–2002, as described by medical subject headings from the National Library of Medicine (NLM) and by citation indexing from the Institute for Scientific Information (ISI).
SOMs show frequently co-occurring terms as nodes that are spatially close. PFNETs show them as nodes with explicit ties. The two kinds of maps will be exemplified here with medical subject headings (MeSH) and cocited authors in a specialty of genetics.
In their extensive review, Borner et al. (5) emphasize that ‘‘painting a big picture’’ is a main goal in domain mapping. This may lead to a strategy of mapping very large co-occurrence matrices in their entirety. Indeed, system designers have made many significant developments in software for such global portrayals of literatures, e.g., THEMESCAPE and VXINSIGHT render literatures as landscapes; GALAXIES and STARRYNIGHT render them as astral bodies (10–12).
In global mapping, system designers present the user with a preformed view, often in 3D, of some sizeable literature. Within the panel of visualization, landscapes invite flyovers; star-fields or other constructs invite flythroughs. In the former, peaks representing major accretions of documents on some subject are likely to exert a powerful pull on the user; in the latter, document points coded as important, e.g., by differences in shape, size, or color, exert a similar pull.
Essentially, the user is engaged in old-fashioned browsing, as of book titles in library stacks, but system designers may minimize or even eliminate labeling of objects in the map because labels clutter precious screen space and block the metaphorical presentation (see examples in ref. 12).
The user explores the view by ‘‘visiting’’ or ‘‘homing in on’’ objects of interest, rather as in video games, but typically cannot remap the literature in pursuit of some new interest because a new map takes hours of computer time to create.
Ours, however, is an alternative way of visualizing knowledge domains, the localized mapping. Perhaps the chief difference is that the localized approach relinquishes scope to increase the user’s control of the mapping process.
... our localized system of mapping more closely resembles online searching. The user starts the process by entering a single term at a web interface. This is consistent with the way most people search the web (13) and is intended to minimize cognitive demands on users.  The system responds to the entry (or ‘‘seed’’) term by forming a list of the terms that co-occur with it, ranked high to low by frequency. The seed term and its 24 next-highest neighbors are then exhibited as a PFNET or a SOM, which the user can switch between.
If the indexing terms used in the mapping are indeed controlled by a formal thesaurus, our SOMs and PFNETs provide an alternative: they display the top listings in what is sometimes called a term’s associative thesaurus (2).
A map of cocited authors is, in effect, an associative thesaurus of authors linked by conjoint use of their works. Again, these linkages may permit useful retrievals that are not otherwise possible (1).
PFNETs and SOMs are dimension-reduction techniques that have been used to visualize the structure of literatures for more
than a decade.
In the context of the movement joining bibliometrics with document retrieval (2, 5, 10), PFNETs have been described by Fowler and colleagues (15–17), McGreevy (18), and Chen (19, 20). Analogous accounts of SOMs have been done by Lin et al. (21), Roussinov and Chen (22), and Chen et al. (23).
The number of links in a PFNET is controlled by two parameters, r and q.
The parameter r, which determines how path weights are computed, is lucidly explained by Fowler et al. (17): ‘‘Path weight, r, is computed according to the Minkowski r-metric. It is the rth root of the sum of each distance raised to the rth power for all links in a path between two nodes. Although the r-metric is continuously variable, simple interpretations exist only for r =1 (path weight is the sum of the link weights in the path), r=2 (path weight is the Euclidean distance), and r=infinity (path weight equals the maximum link weight in the path). One advantage of r=infinity is that one need only assume that the original distance estimates have ordinal properties. Another advantage is that the link structure will be preserved for anymonotonic transformation of the data.’’
The parameter q sets the range within which all paths of length q will be examined in the test of the triangle inequality (24) and removed if they violate it. The larger the value of q, the more extensive the triangle inequality constraint; therefore, links are more likely on a path that violates the rule. If q is one less than the number of nodes, then all of the potential violators are under scrutiny.
The more frequently co-occurring terms, which presumably have greater mutual relevance, occupy more proximate regions on the map. SOMs are designed to render not just the highest co-occurrence counts between terms, but rather relatively high co-occurrences across groups of terms.
They are a softer-focus kind of mapping than PFNETs, but they, too, suggest specific combinations of terms on which the user might want to base retrievals.
This process of self-organization (also known as unsupervised learning) runs over many iterative cycles. In each iteration, the images of term pairs that are strongly related in the high-dimensional space will be moved closer on the lower-dimensional space until stability is reached.
A row from the cooccurrence matrix ‘‘is randomly selected and compared to every output node to determine a winner. Weights of the winning output nodes then are updated so that the next time this input node is presented, this output node will likely be selected again as the winner. In the meantime, nodes surrounding the winning node are similarly adjusted.
The number of iterations needed to train a SOM is often determined empirically (in our case, we optimize the number of training cycles to 2,500).
After the training, input vectors closest in the input space will map to the same regions in the output map. The regions are delineated by areas of nodes in which the elements with the highest value on the vectors are the same.’’
Adjacent areas reflect stronger relationships than nonadjacent areas. Terms in large areas are more influential than terms in small areas.
Slatkin found his own cocited author maps readily interpretable. He was acquainted with every name that appears in Fig. 2. In the PFNET (which he again preferred), he identified the main structural feature, the clusters around himself and Masatoshi Nei, as representing two slightly different subject areas. Both the Nei group and the Slatkin group, he said, have contributed to the literature on genetic flow and population structure, but the Slatkin group has contributed relatively more to the literature on microsatellites (short, repetitive sequences of DNA). Hence, the PFNET was picking up a division he found meaningful.
Interestingly, at the lower left the SOM conjoins Wright, Mayr, and Fisher, who represent the older, pioneering generation in statistical genetics. The SOM algorithm is able to bring this out solely on the basis of their overall cocitation profiles.
If PFNETs seem directive about term relationships, SOMs are merely suggestive. However, their greater ambiguity is perhaps a virtue.
Using AUTHORLINK, the forerunner of PNASLINK, Buzydlowski (9) found that SOMs outperformed PFNETs in capturing the mental models of 20 experts in selected fields of the humanities. ... The experts’ mental models were elicited by having them sort cards bearing authors’ names into intuitively meaningful piles. ... SOMs agreed with the card-sort data better than PFNETs. In the Plato trial, both SOMs and PFNETs were highly correlated with the pooled card-sort data (SOMs, r 0.97; PFNETs, r 0.78), but these correlations were significantly different at P 0.001. In the individual-author trials, a t test of mean agreement scores favored SOMs significantly at P<0.01. 

Huang, Z., Chen, H., Guo, F., Xu, J. J., Wu, S., and Chen, W-H. (2004). Visualizing the expertise space. In Proceedings of the 37th Annual Hawaii International Conference on System Sciences (HICSS'04), IEEE Computer Society.

Huang, Z., Chen, H., Guo, F., Xu, J. J., Wu, S., and Chen, W-H. (2004). Visualizing the expertise space. In Proceedings of the 37th Annual Hawaii International Conference on System Sciences (HICSS'04), IEEE Computer Society.

information visualization/self-organizing map

本論文利用SOM及MDS等資訊檢索與文件處理技術將專家及他們的專長以視覺化的方式呈現在二維圖形上。這個研究的資料是台灣的597位商務與管理方面的學者,針對每位學者提供的研究領域(以國科會的分類,總共包括127個研究領域),在研究時分別對學者及研究領域建立特徵向量,用來產生專家地圖以及專長地圖。學者的特徵向量上每一個成分的二元值代表這位學者是否具有某項研究領域的專長,研究領域的特徵向量上每一個成分的二元值則是代表這項研究領域是否為某位學者的專長。最後將這些資料輸入SOM及MDS進行視覺化,進行MDS處理時兩個特徵向量間的相似程度是以Jaccard模式來進行估算。從結果的專家地圖上,可以發現具有相同與相近專長的學者被映射到相同或鄰近的節點上;在專長地圖上,有共同的理論或分析基礎或是共同的應用範疇的研究領域則會被映射到相同或鄰近節點上,形成群聚。
We focus on a basic form of expertise representation, in which experts are represented by a set of expertise fields. Due to the potential high dimensionality of such expertise data, we chose to examine two dimensionality reduction visualization techniques that have been widely applied in data and document visualization: the Self-organizing Map (SOM) and Multidimensional Scaling (MDS). We present two types of visualization results: the expert map and expertise field map, and provide initial analysis on the effectiveness of these visualizations to support expertise searching and browsing.
One type of set-level document visualization uses interactive scatter plots in different forms, which is also referred to as “dimensions and reference point systems” (Morse, Lewis and Olsen, 2000). Visualization techniques of this type attempt to display additional information about the retrieved documents and to group documents that share the similar characteristics. These characteristics may include the relationship between the documents and the query terms (Ahlberg and Shneiderman, 1994), predefined document attributes such as size, date, source and popularity (Hearst and Karadi, 1997;  Nowell, France, Hix, Heath and Fox, 1996), and user-specified attributes such as predefined topics (Olsen, Korfhage, Sochats, Spring and Williams, 1993).
A second category of techniques attempts to visualize inter-document similarities. This form of visualization is also referred to as “map systems” (Morse, Lewis and Olsen, 2000). There are four major techniques for inter-document similarity visualization: document networks (Thompson and Croft, 1989), physically based modeling techniques (Chalmers and Chitson, 1992), document clustering (Allen, Obry and Littman, 1993; Hearst and Pedersen, 1996) and geographic map metaphors (Chen, Schuffels and Orwig, 1996; Lin, Soergel and Marchionini, 1991).
Mockus and Herbsleb (2002) presented a system named “Expertise Browser” in the context of collaborative software engineering for change management systems. They embedded in some simple visualization elements such as the tree structure and other visual elements to present the expert attributes.
Our research explores this idea by focusing on a simple form of expertise database, where each expert is represented by a list of predefined expertise fields. Each expert can be represented as a binary vector, the elements of which correspond to the fields of expertise and the dimensionality is the number of predefined expertise fields. Each expertise field can also be represented as a binary vector, the elements of which correspond to the experts and the dimensionality is the number of experts in the data set. These representations adopt the vector space model of document representation and share the same high dimensionality characteristic. We chose two commonly used dimensionality reduction techniques for visualizing document space in the literature, the self-organizing map and multidimensional scaling, to generate map metaphors to visualize inter-expert and inter-expertise-field similarities.
The data set we used was resulted from an Internet survey on researchers in business and management fields in Taiwan. The survey was conducted by the National Science Council in Taiwan, and covered almost all the researchers in the business and management field in Taiwan.
The data set contained 597 researchers, who had selected their research interests or expertise from a two-level hierarchy of research fields. ... There were 127 second-level research fields and 2865 researcher-field combinations in the data set. ...  Each of the 597 researchers was represented by a binary vector with 127 elements, which corresponded to the research fields. The expertise similarity between two researchers was derived using vector similarity functions. We also had a dual representation for research fields, similar to the researcher representation. Each of the 127 research fields was represented by a binary vector with 597 elements, which corresponded to the researchers. In this case, similarities among research fields depended on the number of overlapping researchers. Such similarities may reflect the common theoretical/analytical foundations or closely related application domains of the research fields, based on the assumption that researchers typically work on closely related research fields.
The input to MDS is a square, symmetric matrix indicating relationships among a set of objects. Such matrices are usually either similarity or dissimilarity matrices. In the context of our research, a similarity matrix is formed based on the similarity scores of expert/expertise field pairs derived from the Jaccard’s similarity function (Jaccard, 1912).
We conducted a regression analysis to evaluate the general relationship between the researcher similarities and map distances. Researcher similarities were calculated using the Jaccard’s similarity function. A Euclidean distance function was used to calculate the map distances of researcher pairs. ... These statistics showed that SOM and MDS both preserved a large portion of the similarity information, although with certain degrees of distortion.
We observe from Figure 4 (expertise field map) that research fields having underlying similarities based on common theoretical/analytical foundations and application domains were grouped together. ... The expertise field map generated by our visualization techniques revealed meaningful grouping of research fields based on experts’ co-occurrence patterns in multiple research fields.

Moya Anegón, F., Herrero Solana, V. and Jiménez-Contreras, E. (2006). A connectionist and multivariate approach to science maps: the SOM, clustering and MDS applied to library and information science research. Journal of Information Science, 32(1), 63-77.

Moya Anegón, F., Herrero Solana, V. and Jiménez-Contreras, E. (2006). A connectionist and multivariate approach to science maps: the SOM, clustering and MDS applied to library and information science research. Journal of Information Science, 32(1), 63-77.

information visualization/self-organizing map
本研究以1992到1997年間LIS的17種期刊論文為研究資源,抽取論文和作者兩種單位的共被引關係,利用MDS和SOM兩種方法,並配合叢集分析(cluster analysis)產生科學地圖(science map)。作者認為這兩種方法都運用了維度縮減(dimensionality-reduction)的效用,並且具有互補的效果,研究結果發現不管是MDS或是SOM在作者共被引所產生的科學地圖上,大多都可以將圖形上映射的作者分為科學計量學(scientometrics)、引用關係研究(citationist)、書目計量學(bibliometrics)、傳播理論(communication theory)、資訊檢索的認知研究(cognitive information retrieval)和資訊檢索的演算法研究(algorithmic information retrieval),並且都屬於科學研究(science studies)的前三者在圖形上位置相當接近,兩種資訊檢索研究的距離也很相近。在論文共被引上,由於門檻較低,MDS和SOM的科學地圖上,除了原先的科學研究和兩種資訊檢索研究所組成的資訊科學研究以外,還包括圖書館研究與管理學兩個區域。
作者認為MDS產生的科學地圖可以保留論文(或作者)原先在高維度的距離關係,SOM則是保留了它們的型態(topology)關係。
The appearance of studies pertaining to library science reveals the relationship of this realm with information science. Especially significant is the presence of the management on the journal maps.
From a methodological standpoint, meanwhile, we would agree with those authors who consider MDS, the SOM and clustering as complementary methods
that provide representations of the same reality from different analytical points of view. ... This approach may be complemented with other kinds of representation based on network analysis.
Within the techniques of multivariate statistical analysis, three basic methods are included (Egghe and Rousseau, 1990):
1) cluster analysis,
2) principal component analysis (PCA), and
3) multidimensional scaling (MDS).
These methods are referred to as dimensionality-reduction methods because this function is to simplify what might at first appear to be a complex pattern of associations among many entities (Kinnucan, 1987).
SOM is based on the principle of the self-organization and grouping of n-dimensional vectors in a bidimensional space, and has been used to reduce dimensions in a wide variety of document spaces of diverse nature. ... According to Kaski (1997), the SOM presents four important properties for data exploration:
1) Ordered display. The characteristic help us to understand the underlying structures in the series of data.
2) Visualization of clusters. We are able to perceive the clustering density of the different regions of the map.
3) Missing data.
4) Outlier. This enables us to detect unusual cases caused by input errors or similar anomalies.
Journal co-citation mapping is potentially of interest both to the researcher studying the structure of scholarly specialities through the published literature and to the collection manager concerned with developing core journals lists, selecting journals and evaluating collections that serve particular research-oriented constituencies McCain(1991).
The main difference is that the SOM tries to present a locally corrected projection, whereas MDS attempts to preserve all the distances between the points. That is, MDS is distance preserving, while the SOM is topology-preserving.
Tijssen (1993) said that the mental representations of the experts, on an individual micro-level, and the bibliometric maps, on a macro-level, are inherently different.

2013年12月3日 星期二

Rzeszutek, R., Androutsos, D., & Kyan, M. (2010). Self-organizing maps for topic trend discovery. Signal Processing Letters, IEEE, 17(6), 607-610.

Rzeszutek, R., Androutsos, D., & Kyan, M. (2010). Self-organizing maps for topic trend discovery. Signal Processing Letters, IEEE, 17(6), 607-610.

就多筆文件資料以及多個詞語的語料庫而言,可以定義一個詞語-文件的共現矩陣(co-occurance matrix), C,紀錄每個詞語在每筆文件中的出現次數。LDA將矩陣C分解成兩個矩陣Φ和Θ。Φ表示主題-詞語的可能性,其中Φ的第k個向量ϕk中的每個元素即為每個詞語在第k個主題上的出現分布。Θ則是文件-主題可能性,它的第d個向量θd上的元素代表文件d包含各主題的可能性。進一步來說,Θ定義了一個文件空間(document space),這個空間上的每一個維度描述相對應主題在文件上的重要性,當文件彼此在意義上相似的話,他們在文件空間上也會相當接近。如果將研究的文件資料依照發表時間區分成若干的時段,每一個時段tw的主題分布情形θ(tw)可以用這個時段內所有文件的主題分布情形的平均值代表,如下面的式子



在這裡,ND(tw) 是在時段tw內發表的文件數量。過去 [5]便曾經利用LDA對於科學論文進行主題模型的研究,從產生的圖形表現出全球暖化研究在十年間逐漸增加的趨勢。然而有時主題的數量多達100個,若是同時將所有的主題在時間上的變化顯示在圖形上,在分析上便很難觀察每一個主題變化的模式(patterns)。
本研究建議將 自組織映射圖(self-organizing map)的視覺化功能結合LDA模型以追蹤主題在時間上的變化情形,藉由觀察自組織圖上反應(response)模式的變化,可以詳細地了解資料集中內容的改變情形。Kohonen的自組織映射圖 [8]能夠將高維度的資料映射到較低維度的格狀(lattice)圖形上,因此本研究嘗試運用這項技術來解決主題數量龐大時的視覺分析問題。自組織映射圖包括訓練(training)與分群(classification)兩個階段。訓練時重複地隨機選取文件資料,比較每一個節點與文件的相似程度,選取最相似的節點,使選取的節點與鄰近地區的節點都往文件的主題方向調整,訓練階段完成後便能使自組織圖代表全部的文件空間,而每一個節點則能代表主題相似的節點。有別以傳統使用所有資料來訓練自組織圖,本研究僅從文件資料集中隨機抽取500筆文件進行訓練,以減少計算的負荷,並且更大的不同是以KL差異(Kullback-Liebler divergence, KL divergence)取代常見的Euclidean距離(Euclidean distance)做為比較文件主題分布與節點主題分布的相似程度。兩個文件在文件空間上的向量分別是A與B時,它們的KL差異定義為
然而KL差異不符合三角形不等式(triangle inequality)與對稱律,也就是KL(A, B) <> KL(B, A)。為了讓文件間的相似性測量具有對稱性,本研究建議使用下面的式子

分群時將資料集上文件依照發表時間映射到對應時段的自組織圖上,因此每一時段會產生一個自組織圖的分群結果,對圖上的每一個節點nj定義一個分群密度 D(nj)

此處Nj是該節點nj上的文件數量,NmaxNmin分別是該時段分群結果的自組織圖上所有節點的文件最大與最小數量。最後根據這些結果產生二維圖形。針對個別時段的自組織圖分群圖形可以發現該時段內的重要主題,連續觀察所有時段的圖形則可以看出各種主體的變化情形。

本研究以2009年五月到八月ESPN.com與TSN.ca的25754筆體育新聞與評論的RSS feed為分析資料。

In order to track changes in topics over time, simple time-series techniques have been applied to a corpus analyzed using LDA [5]. For instance, in [5], the authors show how scientific papers on global warming gradually increase in popularity over a ten year period. Unfortunately it is not uncommon to have 100+ topics (i.e., dimensions) in a descriptor which makes this analysis difficult for any more than two topics.

Therefore, we propose to use a method similar to the WEBSOM [6] and ProbMap [7] algorithms to perform the trend analysis. These algorithms use Kohonen’s Self Organizing Maps [8] to nonlinearly project a high dimensional feature space onto a low-dimensional output space. Our method merges the idea behind WEBSOM and ProbMap with the work done in [5] to show how a document corpus can change with time.

Therefore, given a corpus, it is possible to define a word-document co-occurance matrix, C, that relates the number of times a given word occurs for a given document. It is desirable to find a way to decompose C because a corpus may contain millions of documents and thousands of words.

Latent Dirichlet Allocation probabilistically decomposes C such that


where Φ and Θ are matrices that express how words, topics and documents are related. Φ contains the topic-word likelihoods or, put another way, how likely any given word is to appear in topic k. Θ is a collection of document-topic likelihoods which relates how likely a document is to contain topic k.

Therefore, for each topic k, there is a vector, ϕk, that contains the word distribution for that topic. Similarly, there exists a vector θd for each document d that contains the topic distribution for that document.

The topic-distribution matrix, Θ, acts as a natural descriptor for all of the documents in the corpus. For any document, d, its associated vector θd  describes its location in a document space.

LDA ensures that documents that are semantically similar (i.e., share many words that are in the same topics) will be close to one another in the document space.

A common choice of divergence in this sort of situation is the Kullback–Liebler (KL) divergence and it is defined for two-dimensional PMFs, A and B as

The KL divergence is not a distance measure (i.e., metric) so it does not satisfy the triangle inequality and KL(A, B) <> KL(B, A). For convenience, it is useful to use a symmetric KL divergence so that the order of the arguments is not important. The symmetric KL measure is defined as


For this paper, a corpus was constructed of 25 754 documents that were collected over a three-month period starting in late May 2009 and ending in late August 2009. The documents were article summaries from RSS feeds from sports news and opinion websites such as ESPN.com and TSN.ca.

The values of the hyperparameters were taken from [5] and NT = 40 since we found that this number of topics provided the best tradeoff between descriptor length and the ability to effectively describe the corpus.

The moving average approach simply produces a document descriptor that is the average descriptor for any particular time window. For each window, tw, we obtain a descriptor, θ(tw), such that



where ND(tw) is the number of documents inside of the time window at time tw. θ(tw) then represents the central tendency of the documents inside of that time window.

Unfortunately, it is very difficult to extend this sort of analysis to more than just two or three topics. Consider Fig. 3. All of the topics are shown on the same plot and it is extremely difficult to visually see any underlying patterns. More importantly, this type of trending only shows the most dominant document types at any point in time.

The Self-Organizing Map (SOM), as proposed by Kohonen [8], maps a high-dimensional feature space onto a lower dimensional representation (usually one or two dimensions).

This allows the map to perform a nonlinear dimensionality reduction on a dataset. However, it can also be used as a classifier, which is what we do in this paper. By examining how many data points each node classifies, it is possible to map complex structures in the feature space onto the lattice.

The first stage trains the SOM on a small, random subset of the original input data. For the dataset used in this paper, that subset is 500 randomly selected documents. This is done since we do not need the SOM to accurately model the entire dataset, just loosely resemble it. This has the added benefit of reducing the computational burden when training the SOM, especially for very large datasets. After training, the SOM will now resemble Θ such that each node is actually the representation of a cluster of similar documents (Fig. 4).

As discussed in Section II-B, KL-divergence is used for the “distance” measure since it well suited to describing the dissimilarity between the document probability vectors.

The second stage filters, or processes, the dataset through the SOM using the sliding window method described in Section III-B. We use the trained SOM as a classifier to determine how many documents in the window are classified by each node. Each node has an associated classification count, Nj, or the number of documents that are associated with node j. A document is associated with a node if nj is the closest node to that document vector.

We define a classification density, D(nj), such that



where Nmax and Nmin are the maximum and minimum count values in the window. This ensures that the map is normalized to be the range of [0,1] so that different windows can be compared.

As the map responses change over time, it defines a 3-D volume (Fig. 7). This volume describes how the map responds over time, as opposed to just observing the response of the SOM at any particular time. As before, the clustering properties of the SOM makes it possible to determine how the document distribution itself changes.