顯示具有 MDS 標籤的文章。 顯示所有文章
顯示具有 MDS 標籤的文章。 顯示所有文章

2015年4月2日 星期四

Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

Wolfram, D., & Zhao, Y. (2014). A comparison of journal similarity across six disciplines using citing discipline analysis. Journal of Informetrics, 8(4), 840-853.

揭露與更好地了解科學傳播 (scholarly communication) 的中研究人員、研究團隊、機構、地區/國家、學科、出版品之間的關係,可以在不同尺度下進行研究。分析資料之間的連結可以從直接引用、共被引、共同著作、詞語或主題的共現,或者隱含式主題等形式進行。期刊間的相似性經常利用共被引,過去曾經進行期刊共被引研究的學科有經濟學 (McCain, 1991)、資訊檢索 (Ding, Chowdhury & Foo, 2000)、資訊系統 (Marion, Wilson & Davis, 2005)、醫療資訊學 (medical informatics, Morris & McCain, 1998)、類神經網路 (neural networks, McCain, 1998)以及半導體研究 (Tsay, Xu & Wu, 2003)。然而共被引研究有以下的困難:首先,在多個學科裡的應用,期刊與期刊間的共被引矩陣可能相當稀疏(Boyack, Klavans, & Börner, 2005)。其次,若是沒有Web of Science等資料庫來源,共被引資料的取得將會相當困難。最後,共被引分析大多只利用共被引次數,而沒有考慮引用來源的任何特性。

資訊計量學的重要研究之一是利用引用資料來確認期刊的學科或專業背景,錯誤的分類結果將會影響期刊在領域內的排名。Glänzel and Schubert (2003)發展一個三個步驟的期刊分類程序,使得期刊裡的文章可以根據參考文獻指定主題。Rafols and Leydesdorff (2009)對Web of Science的主題分類(Subject Categories)和Glänzel and Schubert (2003)的主題分類,比較兩種大型矩陣的分解演算法。Leydesdorff and Rafols (2009) 也使用主題分類引用頻率的引用矩陣研究170多種Web of Science的主題分類之間的關係。Leydesdorff and Schank (2008)以視覺化及動畫的方式呈現期刊之間的關係與它們的跨領域性。

過去的研究有利用期刊引用形象(journal citation image),也就是引用目標期刊的所有期刊的列表,做為期刊之間相似度評估的特徵。在計算上,可將期刊引用形象加入引用的頻率分布做為目標期刊的一種特徵(signature)。然而由於具有影響力與聲譽的期刊可能有相當大量的引用期刊,造成期刊引用形象的計算量較大。Wang and Wolfram (2014)提出使用引用期刊所屬的學科,利用引用學科的引用頻率做為期刊的特徵,來降低計算量。Wang and Wolfram (2014),提出引用學科分析(citing discipline analysis)來評估被引用期刊(cited journals)之間的相似性。引用學科分析根據目標期刊的引用期刊(citing journals)在Web of Science的研究領域上的頻率分布,取代直接。Wang and Wolfram (2014)並將引用學科分析應用於JCR的資訊科學與圖書館學 (Information Science & Library Science)的40種期刊,他們發現在多元尺度與群集分的結果中,若干期刊與其他期刊並不接近,有些同時被歸類於其他學科的期刊並不接近於資訊科學與圖書館學的期刊,而且當這些期刊同時也歸類於較大的相關領域時,其具有較高的影響係數(impact factors)將會降低其他期刊的排名。由於Wang and Wolfram (2014)只探討一個學科內的期刊,無法了解多個學科的期刊是否也具有同樣的情形。

本研究同樣利用引用學科分析估計期刊間的相似性。本研究使用來自於6個學科的120種期刊做為研究資料,其中5個學科彼此間較為接近,包括傳播學 (Communication)、電腦科學-資訊系統 (Computer Science-Information Systems)、教育學與教育研究 (Education & Educational Research)、資訊科學與圖書館學 (Information Science & Library Science)、管理學 (Management),另一個學科-地理學 (Geology) 則較遠。選取期刊的出版期間為1987到2012年,分為三個時期1987–1995、 1996–2004和 2005–2012。利用餘弦測量計算期刊間的相似性,並將估計結果應用於多元尺度(multidimensional scaling)、階層式群集分析(hierarchical cluster analysis)、主成分分析(Principal Component Analysis)等技術。

第一時期可以發現六個學科的相關期刊分為五個群集,其中電腦科學-資訊系統的期刊因為在這時期的數量較少而與資訊科學與圖書館學形成一個群集,地理學的群集與其他群集的距離相當遠。原本主題被分類在傳播學的 Journal of Advertising Research(JAR),在這時期的結果裡與管理學的相關期刊較接近。MIS相關期刊以及 Telecommunications Policy (TP)在資訊科學與圖書館學與管理學之間,Social Science Computer Review (SSCR)則是在資訊科學與圖書館學、傳播學與管理學三個學科間,另外Science Communication (SCOMM)雖然主題被分類在傳播學,但在本研究裡的結果,則是歸入資訊科學與圖書館學的群集。


第二時期電腦科學-資訊系統已經和與資訊科學與圖書館學分開,120種期刊共形成6個群集。在這時期的結果,部分主題分類於資訊科學與圖書館學的期刊被歸入其他學科,包括The Journal of the American Medical Informatics Association (JAMIA)和The International Journal of Geographical Information Science (IJGIS)被歸入電腦科學-資訊系統,後者甚至也很接近地理學;Decision Support Systems (DSS)則是原本在電腦科學-資訊系統分類下,而被歸入資訊科學與圖書館學的群集。Journal of Health Communication (JHC)的主題分類有傳播學和資訊科學與圖書館學,但在這時期更靠近傳播學的相關期刊而被歸類在傳播學的群集內。此外,原本分別在教育學與教育研究以及傳播學的the Academy of Management Learning& Education (AMLE)和JAR都被歸類在管理學的群集裡。

第三時期的六個群集對應到六個學科,但原先主題分類為傳播學的一些期刊更靠近於管理學期刊,而同時具有教育學與教育研究以及資訊科學與圖書館學兩種主題的International Journal of Computer-supported Collaborative Learning (IJCSCL),在結果上明顯可歸類為教育學與教育研究的群集。

有些被指定在某一個學科的期刊結果更接近其他學科,但有些被指定在多個學科的期刊則發現只有接近其中的一個學科。因此這個研究所提出的方法在計算期刊間的接近程度時能夠與期刊共被引分析(journal co-citation analysis)等傳統方法互補。由於以下的幾種原因,有愈來愈多的期刊跨越學科的邊界:1) 期刊出版的範圍愈來愈跨學科(more interdisciplinary),因此吸引愈來愈多其他學科的出版品引用。2) 整合式搜尋工具愈來愈普遍,更容易讓其他學科的作者發現。3) 愈多的期刊加入分析,使得期刊在引用學科的特徵上獨特性減少,彼此間更加相似。

為了更加了解跨學科的期刊,本研究利用主成分分析探討期刊的引用學科特徵。在第一時期,資訊科學與圖書館學的相關期刊中,共有五種期刊屬於兩種成分,包含ISJ, ISR, MISQ等三種管理學期刊以及屬於傳播學和電腦科學-資訊系統的SCOMM和DSS。但在資訊科學與圖書館學主題分類下的ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS以及 MISQ同時也被分類在電腦科學-資訊系統主題內,但卻沒有出現在電腦科學-資訊系統的成份裡。並且JAR和 AMLE分別被分類在傳播學和教育學與教育研究,但在本研究的第一時期結果只屬於管理學。

在第二時期的結果中,IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ等許多MIS期刊同時出現在管理學和資訊科學與圖書館學的成份裡。雖然SSCR僅有被分類在資訊科學與圖書館學主題,但同時包含在資訊科學與圖書館學與傳播學的成分中。另外,雖然IJGIS被分類在資訊科學與圖書館學主題,但並沒有出現在任何成分,顯然是六個學科較周邊的期刊。

第三時期同時在管理學和資訊科學與圖書館學成份裡的MIS期刊更增加了。沒有出現在任何成分的期刊,除了IJGIS以外,還增加了Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM)。

A similarity comparison is made between 120 journals from five allied Web of Science disciplines (Communication, Computer Science-Information Systems, Education & Educational Research, Information Science & Library Science, Management) and a more distant discipline (Geology) across three time periods using a novel method called citing discipline analysis that relies on the frequency distribution of Web of Science Research Areas for citing articles.

Similarities among journals are evaluated using multidimensional scaling with hierarchical cluster analysis and Principal Component Analysis.

The resulting visualizations and groupings reveal clusters that align with the discipline assignments for the journals for four of the six disciplines, but also greater overlaps among some journals for two of the disciplines or categorizations that do not necessarily align with their assigned disciplines.

Some journals categorized into a single given discipline were found to be more closely aligned with other disciplines and some journals assigned to multiple disciplines more closely aligned with only one of the assigned disciplines.

The proposed method offers a complementary way to more traditional methods such as journal co-citation analysis to compare journal similarity using data that are readily available through Web of Science.

Aspects of scholarly communication may be investigated from different levels of granularity to reveal and better understand relationships between researchers, research groups, institutions, regions/nations, specializations/disciplines, publications or publication outlets.

Connections that exist between sources of interest may take the form of direct citations, co-citations, co-authorship, co-occurrence of words or subjects, or more recently, latent topics.

Journal similarity comparison has been frequently studied using co-citations. Journal co-citation studies have been carried out on a number of fields including economics (McCain, 1991), information retrieval (Ding, Chowdhury & Foo, 2000), information systems (Marion, Wilson & Davis, 2005), medical informatics (Morris & McCain, 1998), neural networks (McCain, 1998), and semiconductor research (Tsay, Xu & Wu, 2003). 

One reason co-citation studies tend to focus on individual fields is that the journal–journal co-citation matrix that emerges when multiple disciplines are employed can be quite sparse (Boyack, Klavans, & Börner, 2005).

Co-citation data can also be labor-intensive to extract and are not easily available through citation database sources such Thomson Reuters Web of Science (WoS) without downloading all references from a corpus of articles.

Citation-based data may also be used to identify disciplinary or specialization affiliations for journals. This is particularly important for informetrics studies, where the misclassification of journals may affect the ranking of journals within a given field.

Pudovkin and Garfield (2002) developed a journal relatedness factor based on citing and cited journals. The goal of their proposed method was to help identify thematically related journals.

Similarly, Glänzel and Schubert (2003) developed a three-step process for the categorization of journals that involved pre-defined categories, journal classification and article classification for articles in journals with ambiguous subject assignments based on references.

More recently, Rafols and Leydesdorff (2009) compared the outcomes of two algorithms for the decomposition of large matrices against Web of Science Subject Categories and Glänzel and Schubert’s categorization. The four methods they used resulted in similar map outcomes on a large scale. Leydesdorff and Rafols (2009) also investigated the relationships among 170+ Web of Science Subject Categories using a citation matrix consisting of the subject category citation frequencies. They concluded that a classification scheme could be developed using analytical arguments.

Similarly, Leydesdorff and Schank (2008) visualized and animated the disciplinary ties of three seed journals over time to demonstrate relationships among journals and their interdisciplinarity.

Co-citation analysis relies on citing articles to identify the strength of relationships between the units of interest, whether authors, papers or journals; however, it does not consider any attributes of the source of the citations – only that the citations or co-citations exist. Authors such as White (2001) and Ajiferuke, Lu, and Wolfram (2010) have called for a shift in the focus of citation-based research away from citation counts received by an author of interest to the origin of the citation and its characteristics to assess author impact from a different perspective.

This research investigates the use of data derived from citing journals to assess the similarity of cited journals.

The journal citation image of a target journal, which is determined by the list of journals that cite the target journal, provides an indicator of the reach of a journal. When combined with the frequencies of citation by the citing journals, the frequency distribution of citations provided by the citing journals creates a “signature” for each cited journal. These signatures may be compared using various analytical methods.

One possible challenge associated with using the citing journals themselves to create a signature for a cited journal is the potentially high number of citing journals that an influential and prolific journal might attract.

Wang and Wolfram (forthcoming) proposed a method to reduce the computational overhead associated with the citing journal data. Their method of citing discipline analysis uses the subjects/disciplines assigned to the citing journal and the resulting citation frequencies of the citing disciplines to constitute the cited journal’s signature.

Wang and Wolfram (forthcoming) employed citing discipline analysis to explore journal similarity among 40 high impact journals in Information Science and Library Science (ISLS) as classified in Journal Citation Reports (JCR). They found that some of the journals classified into the ISLS category did not map in close proximity to one another based on multidimensional scaling and cluster analysis. A number of the journals included were also classified into allied fields, but did not cluster or appear in close proximity to a number of journals only classified in ISLS.

The authors noted that how journals are classified can impact journal rankings within a given field, where journals from related, but larger, fields may have higher journal impact factors (IF), which can reduce the rank of journals that are directly in the field. They observed that many of the high impact journals were from allied areas to ISLS.

One limitation of their exploratory study was the focus on a single discipline. Could similar affinities or differences in journal similarity be revealed using citing discipline analysis with journals from multiple fields? Also, by looking at multiple fields, does this pull journals also classified into other disciplines further out of an assigned category when included?

The present research is guided by the following questions:
1. To what extent are high impact journals from allied disciplines similar to one another based on the discipline of the articles that cite a given journal?
2. Do journals classified into multiple disciplines more closely align with one discipline than another or serve as bridges between the disciplines to which they are mapped based on the citing discipline distribution?

The field of Information Science and Library Science was selected as the seed discipline based on its interdisciplinary nature and familiarity to the authors. The top 20 journals based on 2012 JCR impact factors were selected. Four additional allied WoS disciplines were also selected based on the co-classification of journals appearing in the top 20 ISLS list with other disciplines, and the affiliation of information science and library science academic units with other disciplinary units, which demonstrates another type of alliance.

The four JCR disciplines selected comprised:
◦ Communication (COMM) – based on the existence of schools of communication & information.
◦ Computer Science, Information Systems (CSIS) – based on a number of iSchools and journal overlap in JCR.
◦ Education & Educational Research (EDER) – based on a number of ISLS units affiliated with colleges/schools of education.
◦ Management (MGMT) – based on the overlap of journals, particularly in Management Information Systems (MIS).

A sixth, more intellectually distant, discipline, namely Geology (GEOL), was also included. Geology was selected based on the outcomes of the UCSD Map of Science (Börner et al., 2012), where Earth Sciences were mapped as distant from the Social Sciences. By including journals from a more distant discipline, the ability for the citing discipline method to distinguish between more closely aligned and distant disciplines could be tested, where the distinctiveness of allied disciplines may be less defined by including a more distant discipline in the analysis.

A total of 120 journals were studied over the time period 1987–2012. To allow for a comparison over time, the journals were subdivided into three time periods: 1987–1995, 1996–2004, and 2005–2012. 

The data collection method for determining the frequency distribution of citing disciplines used in Wang and Wolfram(forthcoming) was adopted for the present study.

The “Create Citation Report” option in WoS was selected to identify all citing articles. The number of Citing Articles was then selected to retrieve the list of citing articles. The WoS “Analyze Results” feature was next selected for the list of citing articles. On the Results Analysis page, “Research Areas” were selected as the ranking field to provide the tabulated list of citing disciplines.

Salton’s Cosine measure was used determine the similarity between pairs of journals, resulting in a symmetric similarity matrix (Ahlgren, Jarneving, & Rousseau, 2003; Egghe & Leydesdorff, 2009; Leydesdorff, 2006).

Multidimensional scaling (MDS) analysis and hierarchical cluster analysis using SPSS v.20 were employed to visualize and categorize the relationships among the journals for each time period.

To provide a complementary analysis of the hidden groups that may be present in the data, SPSS’s Factor Analysis using Principal Component extraction with varimax rotation was also applied to the data using routines.


Fig. 1 shows the MDS locus of 70 selected journals in the first time period (1987–1995). The raw stress value was 0.01294,and the stress-I was 0.11376.

Only five clusters are shown because a distinctive sixth cluster did not emerge for this time period.

At the five-cluster level of assignment, journals from COMM, EDER, and MGMT form coherent clusters, although based on the MDS map some journals in each field are more closely located to journals in an allied discipline.

The fourth cluster combines journals from ISLS and CSIS. It is possible that the relatively small number of purely CSIS journals for this time period did not provide enough data for these journals to cluster into separate groups.

The MIS journals are situated in the ISLS cluster, but are located between the Library and Information Science (LIS) journals and MGMT journals.

The fifth cluster on the right side of the map consists of GEOL journals and, as would be expected, is quite distinctive from the other five disciplines.

The location of some journals on the map suggests that they are more similar to journals in one of the other given categories in JCR. The Journal of Advertising Research (JAR), for example, is situated with management journals but is only classified with COMM (and with Business, but this discipline is not included in this study).

Some journals classified in two disciplines served as a bridge between the two disciplines in the map. For instance, Science Communication (SCOMM), although classified with COMM journals, is situated in the ISLS cluster, but is in closer proximity to the COMM journals. In the remaining two time periods, SCOMM clusters with the COMM journals. Social Science Computer Review (SSCR), which is classified with ISLS and clusters with the discipline, is situated between ISLS, COMM and/or MGMT journals in each time period. The same is observed with Telecommunications Policy (TP), which bridges ISLS and MGMT for each of the periods of study.

The outcome for the second time period (1996–2004) appears in Fig. 2. The raw stress and stress-I values are 0.01648 and 0.12798, respectively.

In this map, 93 journals were categorized into six clusters, with CSIS separating from the ISLS cluster during this time period.

Of note is the greater number of journals assigned to one or more disciplines but aligning more closely to another discipline or only one of the assigned disciplines.

As an example, the Academy of Management Learning& Education (AMLE) and JAR are situated in the management cluster, and are located relatively far from their assigned disciplines, EDER and COMM, respectively.

Decision Support Systems (DSS) clusters with the ISLS journals but is classified in CSIS only. The same classification is observed for this journal in the third time period.

The Journal of the American Medical Informatics Association (JAMIA) is classified in ISLS and CSIS, but clusters with the CSIS journals for the remaining time periods.

Similarly, the Journal of Health Communication (JHC), which is classified in COMM and ISLS, does not appear to be similar to other ISLS journals and is situated more closely to COMM journals and clusters with them.

The International Journal of Geographical Information Science (IJGIS), which is classified with ISLS journals clusters with CSIS journals for this time period and the third period, although it is at the periphery of the cluster, perhaps indicating the CSIS discipline is the best match of the disciplines studied, but is not a very close match. Proximally, it is situated between CSIS and GEOL journals, which may indicate at least a peripheral similarity to some GEOL journals.

Again, the GEOL journals all cluster together farther from the other disciplinary groups.
Results for the third time period (2005–2012) appear in Fig. 3. The six clusters roughly correspond to the six disciplines. The results of raw stress calculation and the stress-I calculation are still relatively low, at 0.02224 and 0.14914, respectively.

Additional journals classified in COMM map closely to and cluster more closely with MGMT journals.

The International Journal of Computer-supported Collaborative Learning (IJCSCL) is classified with both EDER and ISLS but is clearly situated and clusters with the EDER journals.

Business Strategy and the Environment (BSE), although clustered with MGMT journals, appears to be pulled toward the GEOL journals, indicating a possible relationship with some of these journals.

Once again, the GEOL journals are distinctly clustered away from the remaining journals.

With each time period, more journals cross disciplinary boundaries by clustering with journals from allied disciplines or by mapping more closely to journals in allied disciplines. There may be several influencing factors to account for this observation.

First, the journals indeed may be becoming more interdisciplinary in their publication coverage, thereby attracting more citations from publications in other disciplines.

Second, the journals themselves may not be more interdisciplinary in their coverage, but are now more easily discovered by authors in other disciplines given the wider availability of federated search tools.

Third, with a greater number of journals included in the analysis for each time period, the distinctiveness of the citing discipline signatures may be decreasing, so some journals classified in allied disciplines may appear more similar to one another.

To determine common dimensions from the dataset, a Principal Component Analysis was conducted in SPSS for each time period. Outcomes for the Kaiser–Meyer–Olkin measure of sample adequacy (above 0.7) and Bartlett’s Test of Sphericity (p < .05) indicate the data were appropriate for PCA for all three time periods.


In this period, the six components explain 88.1% of the total variance that correspond to the six disciplines.

There are five journals underlined in the Table 2 that belong to two components, introducing inter-factorial complexity (Van den Besselaar & Heimeriks, 2001; Leydesdorff, 2007), including three MGMT journals (ISJ, ISR, MISQ).

SCOMM and DSS are assigned only to COMM and CSIS by WoS, respectively, but they also appear in the ISLS component, which supports the MDS and clustering outcomes. DSS continues to also load with ISLS for the remaining time periods.

Similarly, JAR and AMLE, journals classified by WoS only in COMM and EDER, respectively, load with the MGMT component for all the time periods in which they appear, but not their classified discipline, lending support for the re-classification of these journals.

The ISLS journals ARIST, EJIS, IM, JASIST, JIS, JIT, JSIS, and MISQ are also classified at CSIS, but do not load with the CSIS component.

Outcomes for 1996–2004 appear in Table 3.

As with the first time period, several MIS journals (IM, ISJ, ISR, JIT, JMIS, JSIS, MISQ) load to both MGMT and ISLS.

SSCR, which is classified with ISLS only, loads into the ISLS component, but also loads with a higher value into the COMM component, perhaps indicating the need for an additional classification assignment.

IJGIS does not load into any of the six components for second or third time period, lending evidence to the peripheral nature of the journal to the six fields studied and indicating it might be misclassified in ISLS.

Component outcomes for 2005–2012 appear in Table 4.

Similar to the previous time period, a growing number of MIS journals load to the ISLS and MGMT components.

As with IJGIS, two other journals, Journal of Chemical Information and Modeling (JCIM) and Journal of Cheminformatics (JCHEM) do not load into any of the six components, indicating a poor association with the six disciplines.

The boundaries between ISLS and CSIS are not as clear in the MDS and cluster analysis outcomes, where combinations of computer science, library and information science and management information systems journals may cluster together depending on the time period. These results may be influenced by the fact that a number of journals in the ISLS area are also categorized in the CSIS or MGMT category, thereby strengthening their relationships.

Despite the influence of the assigned discipline(s) – which then strengthens the disciplinary relationship(s) through journal self-citations, whether or not it is the best fit – some journals appear to be misclassified based on the disciplinary designations of the citations they attract.

As a prime example, JAR, which is only assigned to the COMM field, is situated in the MGMT category for each of the time periods studied for the MDS and clustering analysis as well as with the Principal Component Analysis.

A similar outcome is observed for AMLE for the two time periods in which it is included. It is classified as an EDER journal, but is situated with the MGMT journals based on the analyses conducted.

Several other journals are classified in more than one discipline, but clearly associate only with journals in one of the disciplines. JHC is classified in COMM and ISLS but clusters only with COMM journals. The same is observed for IJCSCL, which is classified in ISLS and EDER but groups only with EDER journals for each of the grouping methods used.

Other journals appear to move between disciplines over time. SCOMM is classified in COMM, but the MDS and clustering outcome places it initially with the ISLS journals for the first time period, and then in the COMM cluster and further away from ISLS in last two time periods.

Some journals, such as JCMC and SSCR, appear to be situated near borders between disciplines, which point to their interdisciplinary appeal or may indicate they serve as bridges between the disciplines.

A number of the ISLS journals that are considered to be library and information science journals (Nisonger & Davis, 2005) appear between CSIS journals and those in the MGMT cluster. In particular, MIS journals appear between MGMT and the ISLS/CSIS cluster for the first time period.

One application of citing discipline analysis that emerges from this analysis is that of decision support for the additional assignment or reassignment of journals to one or more disciplines.

Journal disciplinary classifications should be revisited over time to accommodate shifts in how journals are being cited by other disciplines. ... However, the shifts are at least an indication that the subject affiliations of the citing the journals are changing.

Citing discipline analysis provides, with some modest programming, a relatively easily implemented method for assessing the similarity of journals within disciplines or across allied disciplines that is computationally less expensive than using citing journal-based data.

The analyses reveal distinct groupings of journals based on their disciplinary assignments. As observed earlier in Wang and Wolfram (forthcoming), who only examined journals in a single field, the current research has demonstrated that citing discipline analysis can provide coherent and meaningful disciplinary groupings for journals in allied fields, even when journals from a more intellectually distant field are included.

The clustering and proximity of some journals classified in allied fields has changed over time, perhaps indicating a changing citing relationship between these fields.

2014年8月15日 星期五

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words. Journal of the American Society for Information Science and Technology, 62(10), 1933-1953.

Milojević, S., Sugimoto, C. R., Yan, E., & Ding, Y. (2011). The cognitive structure of library and information science: Analysis of article title words.Journal of the American Society for Information Science and Technology,62(10), 1933-1953.

Scientometrics

圖書資訊學(LIS)為對於記錄下來的資訊(recorded information)和具有文化意義的文物與標本(culturally meaningful artifacts and specimens)有興趣的研究領域(Bates, 2010),包括的領域有檔案學(archival science)、 書目(bibliography)、文獻與文類理論(document and genre theory)、資訊學(informatics)、資訊系統(information systems)、知識管理(knowledge management)、圖書資訊學(LIS)、博物館研究(museum studies)、記錄管理(records management)和資訊的社會研究(social studies of information)。過去有許多研究嘗試定義與描述圖書資訊學的領域並且確認其中包含的研究主題,這些研究使用的方法相當廣泛,包含Järvelin & Vakkari (1990, 1993)採用內容分析(content analysis);Åström (2007, 2010)、Moya-Anegón, Herrero-Solana, & Jiménez-Contreras (2006)和 Persson (1994) 針對期刊或期刊文章進行書目計量分析 (bibliometric analysis) ; Moya-Anegón et al., (2006)和White & McCain (1998)針對作者進行書目計量分析 ;Åström (2002)、 Ding, Chowdhury, & Foo (2001) 和 Janssens, Leta, Glänzel, & De Moor (2006)利用從題名、摘要或全文抽取的詞語進行詞語的共現分析(co-word analysis) ;Sugimoto & McCain (2010)則是用索引詞語的三元共現分析(tri-occurrence analysis) ; van den Besselaar & Heimeriks (2006)利用詞語和參考文獻的組合進行分析;以及Sugimoto, Li, Russell, Finlay, & Ding, (2011)和 Sugimoto & McCain (2010)所使用的主題模型分析方法。

上述的這些方法,許多必須依賴於作者對於領域知識的了解,才能了解領域的主題與認知結構(cognitive structure),例如White & McCain (1998)基於最重要的作家的集群,觀察資訊科學由圍繞在一個微弱中心的許多專業所組成;Åström (2010)則是透過作者與期刊的映射圖說明這個領域的圖書館學(LS)和資訊科學(IS)之間具有差距。除了是認知結構較不直接的指標之外,引用分析另一個的問題是不同的次領域有不同的發表與引用實務。

論文題名包含許多能夠指出該文章內容的詞語(Buxton & Meadows, 1977; Meadows, 1998)。因此,本研究採用的方法是利用期刊論文題名上的重要詞語進行分析。分析的資料來自16種LIS期刊於1988到2007年發表的10344筆論文資料。

選取100個最常出現於題名的詞語。

本研究使用的分析技術包含詞語的相對頻率(relative frequency)並且根據詞語的共現進行叢集,最後並將詞語以及期刊與發表年度等進行多維尺度分析(multidimensional scaling, MDS),產生視覺化的結果。

詞語的共現分析以及階層式集群分析的結果發現三個主要分類LS(圖書館學)、IS(資訊科學)、SCI-BIB(科學計量學-書目計量學)以及兩個較小的分類資訊尋求行為(information-seeking behavior)和書目指導(bibliographic instruction)。LS可再細分為學術圖書館專業(academic librarianship)、公共圖書館專業(public librarianship) (包含館藏建立)、資訊素養和學校圖書館專業(information literacy and school librarianship, technology)、政策(policy)、全球資訊網(the web)、知識管理(knowledge management)、數位圖書館(digital libraries)、電子商務(e-commerce)、法律(law)以及學術出版(scholarly publishing)等主題。IS則包含資訊檢索(information retrieval)、網路搜尋(web search)、分類目錄(catalogs)以及資料庫(database)等主題。SCI-BIB也有書目計量指標(bibliometric indicators)、作者生產力(author productivity)與引用研究(citation study)等主題。整體的結構如下圖

從詞語的使用可以發現LIS中有某些持續出現的核心詞語,但也有一些詞語的使用在20年間有明顯的變化,這些都是與科技相關的(technologically related)詞語,這個現象符合Saracevic(1999)所宣稱的LIS是個科技驅動的(technology driven)領域。大致上來說,LIS內的改變可以從資料庫(database),到數位圖書館(digital libraries),到全球資訊網(the World Wide Web)等詞語使用的移轉上看得出來。

除了科技驅動的特徵外,LIS同時也有很大的範圍在討論資訊尋求行為,這是LS和IS都共同關心的課題。

A number of empirical studies of LIS have been conducted with the aim of describing and defining the field and identifying research areas within it. These studies applied a wide array of approaches: content analysis (Järvelin & Vakkari, 1990, 1993); bibliometric analysis of journals and journal articles (Åström, 2007, 2010; Moya-Anegón, Herrero-Solana, & Jiménez-Contreras, 2006; Persson, 1994); bibliometric analysis of authors (Moya-Anegón et al., 2006,White & McCain, 1998); co-word analysis of both index terms and words extracted from titles, abstracts, and full text (Åström, 2002; Ding, Chowdhury, & Foo, 2001; Janssens, Leta, Glänzel, & De Moor, 2006); tri-occurrence analysis of index terms (Sugimoto & McCain, 2010); analysis of word-reference combinations (van den Besselaar & Heimeriks, 2006); and topic analysis (Sugimoto, Li, Russell, Finlay, & Ding, 2011; Sugimoto & McCain, 2010).

Some notable studies of cognitive structure of LIS have interpreted topics post hoc, by assigning topicality based on knowledge of the author’s domain (e.g., White & McCain, 1998). In White and McCain’s influential visualization of LIS, they concluded that “information science lacks a strong central author, or group of authors, whose work orients the work of others across the board. The field consists of several specialties around a weak center” (p. 343). However, this analysis was based foremost on the clustering of authors, rather than topics. Similarly, Åström (2010) examined the divide between LS and IS components of the field by a bibliometric mapping of authors and journals. Topicality was assigned through expert knowledge of the domains in which these authors wrote and journals published.

Of the various components of textual documents, the titles, and the choice of words in them, are of particular importance. Title words function as “attention triggers” (Bazerman, 1985, 1988). They are devices for capturing interest in the world where information overload is a norm. Title words
have been called “signal-words”1 (Rip & Courtial, 1984) and “macro-actors” or “macro-terms”2 (Callon et al., 1983). Titles of journal articles themselves have undergone a change during the 20th century, becoming more informative, more specific, and containing a larger number of words that indicate article content (Buxton & Meadows, 1977; Meadows, 1998). Leydesdorff (1989) claims that “title words seem to offer a means of making visible the internal cognitive structure” (p. 217) of a discipline. He also claims that “word structure reflects internal intellectual organization in terms
of the codification of word usage in the relevant disciplines” (Leydesdorff, 1989, p. 221). 

Co-word analysis is based on co-occurrence of words (all words, or selected keywords) extracted from titles, abstracts, or text in general, or the index terms assigned by authors or indexers. Co-word analysis is a method that derives “higher level structures from word-occurrence patterns in text” (Chen, 2003, p. 139). Of particular importance in the context of this study is that co-word analysis is “a means to the elucidation of structures of ideas, problems, and so on, represented in appropriate sets of documents” (Whittaker, Courtial, & Law, 1989, p. 473). 

Although co-word analysis has its limitations, (e.g., Leydesdorff, 1997) primarily because of the
change of usage and meaning of words and the lack of context, such analysis has been considered particularly useful in tracking the development of scientific fields over time (Callon et al., 1991; Noyons & van Raan; Rip & Courtial, 1984), which represents another goal of this study.

Although citation analysis is not subject to the same limitation, it is a less direct indicator of cognitive structure. As already mentioned, studies using citations require post hoc assignment of topics. In addition, citation analysis of LIS is less effective in analyzing the cognitive structure of entire fields due to the different publication and citation practices of subfields, thus leaving even large subfields such as LS often invisible.

Selection of journals and articles. Articles from 16 LIS journals were chosen for inclusion in this study. The journals were selected from a ranked list of the most important journals in the field, according to deans and directors of American Library Association (ALA)-accredited, MLS programs in North America (Nisonger & Davis, 2005).

From this journal set, all research and review articles (10,344) published between 1988 and 2007 were included in the analysis.

Identification of the most frequently occurring LIS words and phrases. Word frequency is an important measure in content analysis. This measure is used to identify the most important research topics or concepts in a field by focusing on the most frequently occurring words.

In this study, we base all analyses on the 100 most frequently occurring LIS words or phrases. 

2014年1月25日 星期六

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

Janssens, F., Leta, J., Glänzel, W., & De Moor, B. (2006). Towards mapping library and information science. Information Processing & Management, 42(6), 1614-1642.

本研究利用詞語共現分析(co-word analysis)技術,區分出六個圖書資訊學的研究主題:兩個書目計量學主題、一個資訊檢索主題、一個一般議題、一個網路計量學主題以及一個專利研究主題。

詞語共現分析根據詞語共同在文件出現的現象描述文件的內容,利用共同出現的相對強度呈現領域的概念網絡(concept networks)。目前已經有植物生物學(de Looze and Lemarie, 1997) 、凝態物理(Bhattacharya and Basu, 1998)、化學工程(Peters and van Raan, 1993)、資訊檢索(Ding, Chowdhury, and Foo, 2001)以及 醫學(Onyancha and Ocholla, 2005)等多個領域曾利用詞語共現分析技術來研究領域內的概念網絡。Van Raan and Tijssen (1993)討論基於詞語共現分析的書目計量在知識論的潛力(epistemological potentitals)。相較於共被引分析,詞語共現分析能應用在沒有引用索引的資料,而且共被引分析會因為在領域的變動與趨勢以及引用者的行為而變得複雜(Noyons & van Raan, 1998)。雖然Leydesdorff (1997)認為詞語的意義隨它們與其他詞語關係的頻率及其出現位置,會有所改變;但Courtial (1998)則是認為詞語共現分析中的詞語,並非做為用來代表某種意義的語言單位,而僅僅是文本間的連結指標。

本研究列舉幾個應用文字資訊為基礎的書目計量方法在圖書資訊學研究主題分析的研究:Courtial(1994)以詞語共現分析對這個領域進行探討,結果發現這個領域包含傳統圖書館學、資訊檢索、科學計量學、資訊計量學、專利分析以及最近興起的網路計量學。Glänzel及其同事整合全文為基礎的結構分析(full-text based structural analysis)和傳統的書目計量方法探討書目計量學及其次領域(Glenisson, Glänzel, and Persson, 2005; Glenisson, Glänzel, Janssens, and De Moor, 2005; Janssens, Glenisson, Glänzel, and De Moor, 2005)。

本研究所使用的分析技術包括:文本抽取(text extraction)、前處理(preprocessing)、多維度尺度(multidimensional scaling)以及Ward’s階層叢集(Ward's hierarchical clustering),並且利用向量空間模式(vector space model) (Salton & McGill, 1986)和隱藏語意分析(latent semantic analysis) (Deerwester et al., 1990)測量文件間相似程度的估計值。以論文彼此間的相似程度,將論文映射成二維圖形的結果如下,此圖形並且標示出每篇論文的期刊:

Scientometrics的論文主要分布在標示為1與2的兩個橢圓附近,橢圓1的主題為書目計量,橢圓2則為專利分析。橢圓5上的論文主要來自Information Processing and Management和Journal of the American Society for Information Science and Technology,其主題為資訊檢索。橢圓12的論文傾向於社會方面的主題,除了Journal of the American Society for Information Science and Technology以外,還包括Journal of Information Science和Journal of Documentation。正中央標示為14的橢圓,其主題與網路相關,所有的期刊均有這個主題的相關論文。

以Ward's叢集分析將所有論文進行歸類,最佳的結果共分為六個叢集。本研究並且根據每個叢集上論文的重要詞語以及中心的論文給予叢集的名稱。在二維圖形上標示各種叢集的結果如下:

六個叢集可以圖形上的斜線分為兩群,斜線以下為Bibliometrics1、Bibliometrics2和Patent Analysis,以上則為Webometrics、Information Retrieval和Social Aspects,但六個叢集中以Patent Analysis和其他叢集較分離。書目計量相關論文分為兩個叢集:Bibliometrics1和Bibliometrics2。Bibliometrics1與科學裡的合作關係(collaboration in science)、引用分析(citation analyses)和國家研究成效(national research performance)等主題相關,Bibliometrics2則主要為方法學和書目計量理論相關的論文。

為了找出各期刊分別著重的主題,除了比較上面的兩個圖形,另外還將叢集和期刊的關係映射成圖形。結果發現Information Processing and Management和Information Retrieval幾乎重疊,這個現象表示Information Processing and Management上的論文和Information Retrieval十分相關。Social Aspects和Webometrics相當靠近Journal of the American Society for Information Science and Technology、Journal of Information Science和Journal of Documentation三種期刊。事實上,除了Scientometrics以外,Social Aspects和其他期刊的距離大約相等。最後,Scientometrics則是落在Bibliometrics1、Bibliometrics2和Patent Analysis構成的三角形中心。

The optimum solution for clustering LIS is found for six clusters. The combination of different mapping techniques, applied to the full text of scientific publications, results in a characteristic tripod pattern. Besides two clusters in bibliometrics, one cluster in information retrieval and one containing general issues, webometrics and patent studies are identified as small but emerging clusters within LIS.

The method was developed by Callon, Courtial, Turner, and Brain (1983), more than two decades ago, for purposes of evaluating research. The methodological foundation of co-word analysis is the idea that the co-occurrence of words describes the contents of documents. By measuring the relative intensity of these co-occurrences, simplified representations of a field’s concept networks can be illustrated (Callon, Courtial, & Laville, 1991).

Van Raan and Tijssen (1993) have discussed the ‘‘epistemological’’ potentials of bibliometric mapping based on co-word analysis.

Leydesdorff (1997) analysed 18 full-text articles and sectional differences therein, and considered that the subsumption of similar words under keywords assumes stability in the meanings, but that words can change both in terms of frequencies of relations with other words, and in terms of positional meaning from one text to another. This fluidity was expected to destabilize representations of developments of the sciences on the basis of co-occurrences and co-absences of words.

However, Courtial (1998) replied that words, in co-word analysis, are not used as linguistic items to mean something, but as indicators of links between texts.

Many researchers have used this methodology to investigate concept networks in different fields, among others, de Looze and Lemarie (1997) in plant biology, Bhattacharya and Basu (1998) in condensed matter physics, Peters and van Raan (1993) in chemical engineering, Ding, Chowdhury, and Foo (2001) in information retrieval (IR) and Onyancha and Ocholla (2005) in medicine.

The reason why the emphasis has shifted from co-citation analysis to co-word techniques is twofold. The first reason is a practical one; co-word analysis allows application to non-citation indexes as well. The second relates to methodology; co-citation analysis complicates the combined analysis of field dynamics and trends in the actors’ activity (Noyons & van Raan, 1998).

Bonnevie (2003) has used primary bibliometric indicators to analyse the Journal of Information Science, while He and Spink (2002) compared the distribution of foreign authors in Journal of Documentation and Journal of the American Society for Information Science and Technology.

Bibliometric trends of the journal Scientometrics, another important journal of the field, have been examined by Schubert and Maczelka (1993), Wouters and Leydesdorff (1994), Schoepflin and Glänzel (2001), Schubert (2002), Dutt, Garg, and Bali (2003).

The main journals of the field were also analysed in terms of journal co-citation and keyword analyses (Marshakova, 2003; Marshakova-Shaikevich, 2005).

The co-citation network of highly cited authors active in the field of IR was studied by Ding, Chowdhury, and Foo (1999).

Finally, Persson (2000, 2001) analysed author co-citation networks on basis of documents published in the journal Scientometrics.

Courtial (1994) has studied the dynamics of the field by analysing the co-occurrence of words in titles and abstracts. Courtial described scientometrics as a hybrid field consisting of invisible colleges, conditioned by demands on the part of scientific research and end-users. Although this situation might have somewhat changed during the last decade, this conclusion illustrates how heterogeneous the much broader field of LIS – comprising subdisciplines such as traditional library science, IR, scientometrics, informetrics, patent analyses and most recently the emerging specialty of webometrics – nowadays is.

In recent papers, Glenisson, Gla¨nzel, and Persson (2005), Glenisson, Gla¨nzel, Janssens, and De Moor (2005), Janssens, Glenisson, Gla¨nzel, and De Moor (2005) have applied full-text based structural analysis in combination with ‘‘traditional’’ bibliometric methods to bibliometrics and its subdisciplines.

The full-text analysis consisted of text extraction, preprocessing, multidimensional scaling, and Ward’s hierarchical clustering (Jain & Dubes, 1988).

In short, the textual information is encoded in the vector space model using the TF-IDF weighting scheme, and similarities are calculated as the cosine of the angle between the vector representations of two items (see Salton & McGill, 1986; Baeza-Yates & Ribeiro-Neto, 1999).

The term-by-document matrix A is again transformed into a latent semantic index Ak (LSI), an approximation of A, but with rank k much lower than the term or document dimension of A. A latent semantic analysis is advisable, especially when dealing with full-text documents in which a lot of noise is observed.

One advantage of LSI is the fact that synonyms or different term combinations describing the same concept are mapped on the same factor, based on the common context in which they generally appear (Berry et al., 1995; Deerwester et al., 1990).

A lot of time was devoted to the detection of phrases. Since the best phrase candidates can be found in noun phrases, the programs LT POS and LT CHUNK4 have first been applied to detect all noun phrases in the complete document collection.

MDS represents all high-dimensional points (documents) in a two- or three-dimensional space in a way that the pairwise distances between points approximate the original high-dimensional distances as precisely as possible (see Mardia, Kent, & Bibby, 1979).

The agglomerative hierarchical cluster algorithm using Ward’s method (see Jain & Dubes, 1988) was chosen to subdivide the documents into clusters. ... One of the disadvantages of agglomerative hierarchical clustering is that wrong choices (merges) that are made by the algorithm in an early stage can never be repaired (Kaufman & Rousseeuw, 1990). What we sometimes observe when using hierarchical clustering is the forming of one very big cluster and a few small very specific clusters.

The journal Scientometrics can be largely separated from the other journals (which is also confirmed by the different term profile in the table of Appendix 1), and exhibits two different foci (best visible in Fig. 4).



The first ‘‘leg’’, indicated by the ellipse with number 1 and by and large containing the first focus of the journal Scientometrics, clearly contains papers in bibliometrics. The 10 best TF-IDF terms for ‘‘leg’’ #1 are: citat, cite, impact factor, self citat, co citat, scienc citat index, citat rate, isi, countri and bibliometr.

The second ‘‘leg of Scientometrics’’, indicated by number 2, is characterised by the best terms patent, industri, biotechnolog, inventor, invent, compani, firm, thin film, brazilian and citat. The JIS paper (#3) embedded in this patent ‘‘leg’’ might be considered an outlier for that journal, but it was put in the right place since it is concerned with ‘‘The many applications of patent analysis’’ (Appendix 2: Breitzman & Mogee, 2002).

An important focus of LIS is indicated by ellipse #5 and can be profiled as ‘‘Information Retrieval’’ (IR) when looking at the highest scoring terms: queri, search engin, web, node, music, imag, xml, vector and weight.

The fourth distinguishable subpart of LIS (#12) is about digit, internet, servic, seek, behaviour, health, knowledg manag, organiz, social and respond; so encompassing the more social aspects.

The remaining large subpart is somewhat the central part (#14). It consists of papers leading to a mean profile containing the terms web, web site, classif, domain, web page, languag, scientist, region, catalog, and web impact factor.

The term network of Cluster 1 allowed the conclusion that the papers belonging to this cluster are concerned with domain studies, studies of collaboration in science, citation analyses, national research performance and similar issues.



The medoid is a paper by Persson et al. on ‘‘Inflationary bibliometric values: The role of scientific collaboration and the need for relative indicators in evaluative studies’’ (Appendix 2: Persson et al., 2004). This is a methodological paper with strong implications for research evaluation, combining research collaboration with citation analysis and construction of national science indicators.

The smaller bibliometrics cluster (Cluster 3: manually labelled as ‘‘Bibliometrics2’’) is of more methodological/theoretical nature.




The medoid is the state-of-the-art report ‘‘Journal impact measures in bibliometric research’’ (Appendix 2: Gla¨nzel & Moed, 2002).

The term networks for the two bibliometrics clusters just described contain a few overlapping terms (bibliometr, chemistri, citat, citat rate, cite, cluster, countri, impact factor, isi, physic, rank and scienc citat index). The MDS plot of Fig. 15 confirms that there is no clear border between Bibliometrics1 and Bibliometrics2, but that there is a gradual transition.

The almost tiny Cluster 2 (19 papers, Fig. 10) represents patent analysis.


A paper on ‘‘Methods for using patents in cross-country comparisons’’ forms the medoid of this cluster (Appendix 2: Archambault, 2002).

Cluster 4, with 282 papers, is the largest one. We have labelled it ‘‘Information Retrieval’’.


The medoid paper is entitled ‘‘Querying and ranking XML documents’’ (Appendix 2: Schlieder & Meuss, 2002).

Cluster 5, with 62 papers, belongs to the small clusters. Both terms and papers close to the medoid characterise this cluster as ‘‘Webometrics’’.


The medoid paper is entitled ‘‘Motivations for academic web site interlinking: evidence for the Web as a novel source of information on informal scholarly communication’’ (Appendix 2: Wilkinson et al., 2003).

Cluster 6 (213 papers) proved to be the most heterogeneous cluster. We have labelled it ‘‘Social’’, however, we could also have called it ‘‘General & miscellaneous issues’’.



‘‘Approaches to user-based studies in information seeking and retrieval: a Sheffield perspective’’ is the title of the medoid paper (Appendix 2: Beaulieu, 2003).


The Patent cluster can be clearly separated from the rest of LIS. The subspace under the line is almost completely occupied by Bilbiometrics1, Bibliometrics2 and Patent.




IR and IPM almost collide in this 2D projection (Fig. 20). This means that Cluster 4 (‘‘IR’’) is very close to the scope of this journal.

The ‘‘Social’’ cluster with general and miscellaneous topics as well as ‘‘Webometrics’’ are close to JIS, JDoc and JASIST, too. Moreover, the ‘‘Social’’ cluster is almost equidistant to all traditional journals in Information Science.

The remaining three clusters, namely Bibliometrics1, Bibliometrics2 and Patent, form a triangle in the centre of which the journal Scientometrics is located. The relatively large distances among these clusters and between each cluster and the journal, strongly indicate that a quite large spectrum of bibliometric, technometric and informetric research using different vocabularies is covered by the journal Scientometrics. This observation is in line with the findings by Schoepflin and Gla¨nzel (2001) that scientometrics consists of several subdisciplines such as informetric theory, empirical studies, indicator engineering, methodological studies, sociological approach and science policy; and that case studies and methodology became dominant by the late 1990s. At the end of the 1990s, also technology related studies based on patent statistics became an emerging subdiscipline of the field.

We have found two clusters in bibliometrics, of which a big one in applied bibliometrics/research evaluation and a smaller one in methodological/theoretical issues; also we have found two large clusters in information retrieval and general and miscellaneous issues and, finally, two small emerging clusters in webometrics and patent and technology studies. Within the IR cluster, we have found a small subcluster on music retrieval, which might be a temporary phenomenon since the journal JASIST has published a special issue on this topic.

According to the expectation, IR, General issues and Webometrics were represented by four of the five journals, namely JIS, IPM, JASIST and JDoc, while the two bibliometrics and the patent clusters were the domain of the journal Scientometrics.

2014年1月24日 星期五

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

Lu, K., & Wolfram, D. (2010). Geographic characteristics of the growth of informetrics literature 1987–2008. Journal of Informetrics, 4(4), 591-601.

本研究探討在地理上的生產力遷移(shifts in productivity)是否發生在書目計量學(bibliometrics)、資訊計量學(informetrics)和科學計量學(scientometrics)等計量學(metrics)領域,也就是歐洲的貢獻明顯地成長,並且北美的貢獻相對來說有減少的情形。

有關計量學的研究,Hood and Wilson (2001)和Stock and Weber(2006)等研究都分析了這個領域的文獻成長情形。Hood and Wilson (2001)回顧了計量學領域的發展,並且比較bibliometrics、scientometrics和informetrics的相關文獻,發現bibliometrics還是在相關領域上使用最廣泛的詞語。Stock and Weber(2006)從觀察中確認這個領域從1980年後便持續地成長。Wolfram (2008)則發現在計量學領域中,北美的文獻有明顯地減少而歐洲則是急遽地增加的情形。

本研究利用bibliometrics、scientometrics、informetrics、cybermetrics、webometrics、citation analysis、link analysis和citation indexes做為檢索的問句,同時再加上Scientometrics和Journal of Informetrics兩種期刊的論文,從Web of Science資料庫中進行檢索。結果共檢索出4404筆論文資料。

在這些論文資料裡,共有75個國家。以地區來區分,歐洲在每個時段上具有最大的貢獻,不論是數量或所占比率都有成長,亞洲所佔的相對比例在22年間有很大的成長,北美雖然在數量上有成長,可是相對的比例呈現緩慢的下降。每個地區的作者會偏好在本身地區的期刊上發表,舉例而言,歐洲作者發表論文的前五個期刊中有四個歐洲期刊,南美也有類似的情形,但是亞洲的情形例外,前五個期刊中有四個是歐洲期刊,另一個則是北美的期刊。

自1990年代中期後,國家間的合作情形增加許多,之前國際合作的論文每年為1到19篇,2008年已大幅增加為96篇。美國是國際合作佔最多的國家,但以地區來說,歐洲平均每個國家的國際合作數為5.78篇論文,多於世界其他部分的4.47篇論文。

此外,歐洲則有許多具有國際合作經驗的機構,共有16所研究機構有國際合作經驗,北美則有8所,亞洲有1所。機構間的合作來說,在1987年每篇論文平均只有1.1個機構,但在2007年則增加為1.96。

本研究且利用MDS、VOSviewer和Pajek將這些論文上的國家與機構之間的合作關係,呈現為圖形。

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).


This investigation was prompted by interest in whether shifts in productivity based on geography are observed in the bibliometrics, informetrics and scientometrics areas.

One of the authors conducted a pilot study to determine whether there have been clear declines in North American contributions to the metrics literature base (Wolfram, 2008). The author found that there was indeed a notable relative decline in North American contributions and a sharp increase in European contributions.

Hood and Wilson (2001) examined the growth of literature of the metrics area. They provided an historical treatment of the development of these areas that included earlier studies of the field. In their research, literature associated with bibliometrics, informetrics and scientometrics was compared for the period 1968–2000. The authors noted that bibliometrics was still the most widely used term for metrics research.

More recently, Stock and Weber(2006) conducted a Web of Science search for records specifically including metrics terms and allied areas. They observed contributions had grown substantially since 1980.

Search parameters included the Boolean ORed result of bibliometrics, scientometrics, informetrics, cybermetrics and webometrics, in truncated form (e.g., webometri*), along with the phrases “citation analysis”, “link analysis” and “citation indexes”. ... These search results were ORed with the two primary journals that publish metrics research that are indexed by WoS, namely Scientometrics and the Journal of Informetrics.

A pair-wise comparison of all collaborations at the national and institutional levels was then conducted from which a cooccurrence matrix could be compiled.

Multidimensional scaling (MDS) analysis was used to visualize the relationships among countries. Because the data represent a type of similarity measure represented as a symmetric matrix, SPSS PROXSCAL was used to construct the map, as recommended by Leydesdorff and Vaughan (2006).

The recently developed visualization tool VOSviewer (van Eck &Waltman, 2010) was also used to provide an alternate visualization of the relationship outcomes. Like MDS, VOSviewer (http://www.vosviewer.com/) relies on a distance-based approach to mapping informetric relationships. Instead of using more traditional similarity measures to produce a normalized outcome for co-occurrences as used in MDS, relationships are based on association strengths, so the algorithm is somewhat different than PROXSCAL and, therefore, can produce different outcomes. Details of the comparison of different measures can be found in van Eck and Waltman (2009).

The network visualization software Pajek (http://vlado.fmf.uni-lj.si/pub/networks/pajek/) was used as well. Unlike the distance-based mapping of PROXSCAL and VOSviewer, Pajek produces directed or undirected network maps, with the strength of the relationships represented by the thickness of connecting lines between vertices on the map. Distances are used more for clarification, but proximities do not necessarily indicate a stronger relationship.

The search parameters retrieved 4404 publications.

Europe shows the highest levels of contribution, both in absolute and relative terms over the time period of the study. Growth patterns in absolute terms are nonlinear based on trend line analysis in MS Excel; however, the R-squared goodness-of-fit values for even the best fitting models (higher order polynomials) were never more than 0.95, indicating a less than desirable fit.

Relative contributions based on geographic divisions have been largely stable. An exception is Asia, which had an increasing relative contribution over the 22-year time frame of the study. Although North American contributions have continued to increase in absolute numbers, the relative contribution shows a slow average decline over time.

The top five journals listed for each continent demonstrated a regional preference for publication outlets from that region. So, for example, four of the top five journals for European publications were published in Europe, and four of the top five journal outlets for South America were South American. The exception to this was Asia. Four of the top five journals for Asian publications were European and one was North American. This outcome may be a reflection of the data extraction method, the indexing practices of WoS, or a preference during the study time frame for Asian scholars to publish in Western journals.

Seventy-five countries were represented in the record set.

The number of metrics papers published annually that represent collaborations between two or more countries has increased greatly since the mid-1990s. Prior to this time, the number of internationally collaborative papers ranged from 1 to 19 papers annually. Over the last decade this number has increased to a high of 96 papers in 2008.

In metrics research, the United States also has the highest share of international collaborations, but the average number of collaborations with European countries was higher (5.78 publications per country) than for other parts of the world (4.47 publications per country).

Sixteen of the institutions on the list are European, eight are North American, and one is Asian. The United States has the largest number of institutions represented (five), followed by Belgium (four – note: one institution merged with another institution to form a new entity).

There has been steady growth in inter-institutional collaboration over the 22 years. The mean number of collaborative institutional partners within the dataset has steadily increased from a low mean of 1.1 institutions per publication in 1987 to a high of 1.96 institutions per publication in 2007.

Europe, and in particular Western Europe, clearly dominates in the production of metrics literature. The United States continues to be the largest singular contributor, but this appears to be changing. North American contributions as a whole continue to increase, but represent a smaller percentage of worldwide production. European contributions have grown tremendously, especially during the last 5 years of the study period. This same period is marked by impressive growth from Asia.

It should be noted that WoS increased its coverage in 2008 by including more regional journals. These inclusions possibly could contribute to the increase in Asian contributions, but the observed growth for Asia was already evident prior to any such additions.

International and inter-institutional collaborations do not necessarily reveal strong geographic affinities, although the multiple institutional affiliations by a number of scholars associated with Flemish institutions do contribute to the strengthening of regional ties. Undoubtedly, the growth of the Internet and increasing availability of other telecommunication technologies have made these collaborations less distance dependent.

2013年12月19日 星期四

White, H. D. and McCain, K. W. (1998). Visualizing a discipline: An author co-citation analysis of Information Science, 1972–1995. Journal of the American Society for Information Science, 49, 327-355.

White, H. D. and McCain, K. W. (1998). Visualizing a discipline: An author co-citation analysis of Information Science, 1972–1995. Journal of the American Society for Information Science, 49, 327-355.

vis_paper

本論文探討作者共被引方法,並將其應用在資訊科學。這個研究分析了1972到1995年間12份資訊科學相關期刊內的作者共被引資料,以每八年為一期,所以整個24年研究共3期,每一期均找出被引用次數最多的前100位作者,整個期間共120位,其中的75位在三個時間都有出現。本研究使用的方法與結果分別如下
1) 對120位作者與其他作者的共被引次數形成的矩陣進行Pearson相關係數分析,再利用主成分分析(principal components analysis)與最大變異轉軸(varimax rotation)進行因素分析(factor analysis),了解資訊科學的專業(specialty)結構。以特徵值(eigenvalue)大於1決定抽取的因素數目,每一個因素代表一個專業,如果作者在某一特定的因素上具有0.3以上的負荷(loading),便視為引用者一般認為這位作者具有這個專業。由於作者可能在多個因素上都有超過0.3的負荷,因此每位作者可能會具有多種專業。在本研究中,共抽取出12個因素,可以解釋84%的變異情形,這些因素中前8個特徵值較大,可以從作者辨識的資訊科學專業為 a)設計與評估文件檢索系統的實驗檢索(experimental retrieval);b) 研究科學研究文獻關連的引用分析(citation analysis);c)應用於實際資料庫的實務檢索(practical retrieval);d) 從文字及書目資料分布規律探討數學模型的書目計量學(bibliometrics);e) 研究圖書館自動化、圖書館運作等議題的一般圖書館系統理論(general library systems theory);f) 研究資訊需求與使用的使用者理論(user theory);g) 研究科學的社會系統(social system of science)的科學傳播(scientific communication);h)OPAC ;另外幾個因素則由研究被引入資訊科學的其他領域學者組成。根據各專業上的作者交互情形以及下述映射圖的結果,資訊科學可以分為對於知識文獻以及其社會脈絡的分析研究和人-電腦-文獻的介面研究等兩個次學科。
2) 根據120位作者在3個時期的平均共被引次數,分析他們在各時期的代表性與影響力。
3) 以作者的共被引次數矩陣所產生的相關係數,也就是他們被引用者一般認定的相似性,做為他們之間的關連性,利用多維縮放技術ALSCAL,將每個時期前100位作者映射成圖形,使得共被引次數分布彼此相似的作者在產生圖形上的映射點有較近的距離。並以叢集分析技術CLUSTER進行完全連結叢集(complete linkage clustering),將作者根據他們之間的關連性分為次學科。結果發現,屬於同一個專業的作者在圖形上的映射點彼此間的距離比較近。並且如先前類似的研究所指出的,資訊科學很明顯地可以區分為資訊檢索及領域分析(domain analysis)等兩個次學科。比較不同時期的圖形,雖然少部分的作者映射點有明顯移動,但大多數的作者其映射點的位置相當穩定。
4) 從三個時期的映射圖上作者映射點位置的改變情形產生映射圖,表示作者引用形象(citation image)的改變。
5) 以經典作者(canonical auhtors)在三個時期的共被引相關係數為輸入,利用INDSCAL評估三個時期維度的重要性,從引用的角度驗證學科是否發生典範轉移的情形。結果發現表示「人-電腦-文獻」介面(human-computer-literatures interface)的第二個維度比起表示資訊科學主題專業的第一個維度在三個時期的重要性有大的變化,1972-1979年的第一時期這個維度的重要性不高,1980-1987年的第二時期其重要性則大幅增加,到了1988-1995年第三時期則稍微減少。許多研究者認為資訊科學在1980年代有典範轉移(paradigm shifting)發生,White and McCain上述的結果可以驗證這個現象。

We defined the authors of information science as all those cited in 12 journals, as listed below. Authors were ranked in order of citedness for the entire period covered by Social Scisearch, 1972–1995. Co-citation data were retrieved for all pairs in the top-ranked 120, from which we produced:
1) A factor analysis of the 120 authors for the entire 24-year span, 1972–1995, which reveals the specialty structure of the discipline. Factor analysis, unlike multi-dimensional scaling and clustering, can show an author’s contribution to more than one specialty.
2) Analyses of the 120 authors’ mean co-citation counts, which indicate their standing and influence in the discipline as of 1972–1979, 1980–1987, 1988–1995, and at the end of the three periods combined.
3) Two-dimensional maps of the top 100 authors in each of the 8-year periods (made with ALSCAL, the SPSS multidimensional scaling program) .
4) A map of authors whose ‘‘citation images’’ changed markedly over the years of our study.
5) A two-dimensional composite map of the authors who are in the top 100 in all three periods—some 75 in all. Their most cited works arguably make up the canonical literature of information science. Certain statistics generated by the mapping routine (INDSCAL, a part of ALSCAL) may bear on paradigm shift in the discipline.

In any field of scholarship, writers make judgments as to who has written on what, using what methods, and they reflect the judgments in their citing practices. Aggregated over time, these practices assume definite structure: Writers show commonalities in how they judge the subject matter, methodology, and intellectual style of other writers; for example, they often attach the same meanings and significance to precedent works (Cozzens, 1985; Small, 1978) .

It suggests how authors are commonly viewed on two dimensions, often interpretable as subject matter and style of work. ... Author clusters placed on these two dimensions can be interpreted as specialties within a discipline (White, 1990a, 1990b) .

What is actually mapped is an author’s citation image. Everyone ever cited has one, but only those who have been cited in many writings are likely to figure in ACA. In the latter case, the image has a constant part, the author’s identity as it is rendered in successive reference lists. The image also has a variable part, the gradually increasing set of other author-names that co-occur with a given author in those lists. At the end of a time period, ACA sums up the record by mapping the author as a single point among other selected author-points on the basis of the repeated co-occurrences. Authors with similar profiles of co-occurrences are displayed close together.

The decisive argument for ACA is that it enables one to see a literature-based counterpart of one’s own overview of a discipline.

As is well known, the closeness of author points on such maps is algorithmically related totheir similarity as perceived by citers. We use Pearson r as a measure of similarity between author pairs, because it registers the likeness in shape of their co-citation count profiles over all other authors in the set.

The raw co-citation counts were converted to Pearson r correlation matrices by the FACTOR routine in SPSS, and factors were extracted by principal components analysis with varimax rotation. The default criterion of ‘‘eigenvalues greater than one’’ determined the number of factors extracted.

The Pearson r correlation matrices for ALSCAL and CLUSTER in SPSS were generated with another SPSS rountine, CORRELATIONS ( cf. McCain, 1990) . They were treated as nonmetric (ordinal) similarity data in ALSCAL and grouped by the complete linkage method in CLUSTER. Subdisciplinary groupings of the author points on the maps are based on the dendograms from CLUSTER.

Authors in the top 100 in all three periods—‘‘the canonical 75’’—were separately mapped with INDSCAL, a routine in the ALSCAL bundle that does a specialized kind of multidimensional scaling. The input data to INDSCAL are judgments on the similarity of a set of stimuli by a set of judges. INDSCAL reveals not only the judges’ composite view of the stimuli in multidimensional space, but the weight each individual judge gives each dimension; INDSCAL is short for ‘‘individual differences scaling.’’ We used the individual weights in a new way to explore the notion of ‘‘paradigm shift’’ as it affects the canonical 75.

The two-dimensional space in which the authors appear is relative, not absolute, and it fails to capture certain relationships among oeuvres that appear in higher dimensionality.

Specialties
The results of the factor analysis, incorporating 24 years’ worth of data for the 120 authors, are presented in Table 3. ... Twelve factors were extracted; jointly (R2 ) , they explain 84% of the variance. ... The first eight factors alone explain 78% of the variance. All have seven or more authors with loadings greater than 0.60 and may be interpreted as specialties within the discipline.

The two biggest specialties, obviously, are experimental retrieval, which focuses on the design and evaluation of document retrieval systems, and citation analysis, which focuses on the interconnectedness of scientific and scholarly literatures, usually with data from ISI.

The third biggest specialty we have labeled practical retrieval. Unlike the experimental retrievalists, the authors in this group, rather than working with content-neutral indexing theory, thought experiments, or document testbeds, have tended to discuss retrieval in terms of ‘‘real world’’ databases; terms such as ‘‘INSPEC’’ or ‘‘DIALOG’’ occasionally profane their pens.

The next specialty we call bibliometrics—a word often used to subsume the specialty we labeled citation analysis. However, unlike the citationists, the authors who load primarily here, including the pioneers Lotka, Bradford, and Zipf, are most interested in mathematically modeling certain regularities in textual or bibliographic statistical distributions, irrespective of the literatures from which they come.

General library systems theory is a not altogether satisfactory name for a body of writings on library automation, library operations research, library and information service policy, retrieval system evaluation, and many other interconnected topics.

The specialty we call user theory is appropriately headed by Dervin, author of a highly cited chapter on ‘‘information needs and uses’’ in the 1986 ARIST. ... It will be seen that authors who write about literatures—the citationists, bibliometricians, and scientific communication people—never load above 0.30 on this factor, apparently because citers do not perceive their work as having the right psychological content. On the other hand, quite a few retrievalists load above 0.30, and this suggests the nature of the cognition involved. It has to do with problem-solving at the interface where literatures are winnowed down for users with: Question formulation, search strategies, information-seeking styles, relevance judgments, and the like.

Authors loading mainly on scientific communication all have strong disciplinary identities outside L&IS—for example, in sociology. They may be thought of as explicating the social systems of science, including those in which formal publication of results is an important (but not the only important) part. The sociologists among them all have loadings, some quite high, in citation analysis, confirming their relevance to the study of scientific literatures.

The design of computerized library catalogs, especially for subject searching, is the province of authors who load on OPACs (online public access catalogs) . It makes sense that leading authors here, such as Matthews, Hildreth, Cochrane, and Drabenstott, load secondarily in practical retrieval, just as several of the primary authors there, such as Borgman and Fidel, also turn up here.

As was said, the chief remaining factor seems a collection of authors in other disciplines from whom information science has imported ideas—e.g., cognitive science (Winograd) , information theory (Shannon) , computer science (Knuth)—that are all variously relevant to the central concern of information science, the human–computer–literature interface.

In fact, as both author cross-loadings and the maps below suggest, almost all of the factors or specialties in Table 3 can be aggregated upward into two larger subdisciplines: (1) The analytical study of learned literatures and their social contexts, comprising citation analysis and citation theory, bibliometrics, and communication in science and R&D; and (2) the study of the human–computer–literature interface, comprising experimental and practical retrieval, general library systems theory, user theory, OPACs, and indexing theory.

The Maps
Figures 2 through 4 are our 8-year period maps. We shall use them to explore the idea, introduced earlier, of two subdisciplines in information science.We operationalize this idea as the last two clusters joined in a complete-linkage clustering of 100 authors. These final clusters, which are brought together only after all closer ties have been exhausted, are separated by an angled line superimposed on each map.

We have not, as in the past, drawn lines around smaller clusters of authors corresponding to their specialties. The crowding of many names on the maps makes this difficult, and, besides, the specialties are better conveyed by the factor analysis of the earlier section. To a great extent, however, the authors forming specialties in the factor analysis will be found to have been placed near each other in the maps.

The first finding to note is the overall stability of information science, as here defined. Some author-points undergo remarkable changes of position from map to map, but many more authors stay put in discernible specialties. Fully 75, moreover, persist through all three maps.

We conclude that author co-citation analysis is useful for rendering the inertia of fields. In other words, it objectively captures the slow-changing divisions on which one’s subjective sense of ‘‘semi-permanent’’ disciplinary structure rests.
Co-citation analysis of papers, as opposed to authors, captures disciplinary history at a different, faster rate, which may better suit fields with livelier research fronts than information science.

However, ‘‘domain analysis,’’ as put forward by Hjørland and Albrechtsen (1995) , seems a more appropriate choice. It incorporates citation analysis and bibliometrics, but also a range of topics broader than what ‘‘bibliometrics’’ usually implies— for example, scholarly and professional communication, parts of sociology of science and sociology of knowledge, interdisciplinary linkages, discourse communities, and disciplinary vocabularies (cf. Beghtol, 1995) .
ACA’s confirmation of expert judgments by Hjørland and Albrechtsen, Persson, and the Vickerys is consistent with the claim that citation databases can be exploited for non-experts in a form of AI.

The axes in INDSCAL maps are not subject to rotation and are supposed to be maximally interpretable. Thus prompted, we think the horizontal axis conveys, as in past studies, the range of subject specialties within the subdisciplines of domain analysis and information retrieval. ... Coherent groups from left include the citationists, the arc of bibliometricians across the top and the philosophically orienting figures across the bottom, ‘‘generalist’’ writers such as Smith, Wilson, Saracevic, and Swanson, and the hard and soft retrievalists. The plot generally makes good sense. For example, it is easy to accept Bookstein, Tague-Sutcliffe, Kantor, Buckland, Vickery, and Shaw as transitional figures between the retrievalists and the bibliometricians.
The more interesting vertical axis reflects another subject-related continuum. Information science deals, we said earlier, with ‘‘the human–computer–literature interface.’’ If so, then the top pole represents a relative emphasis on literatures as objects of study, and the bottom, a relative emphasis on people or users. The same polarity can be inferred in earlier maps. Figure 4 showed that when a literature theoretician like Egghe enters, it is automatically at the top, whereas a user theoretician like Dervin is automatically placed at the bottom.

However, INDSCAL is expressly designed to reveal differences in the importance of each dimension to whoever is judging the similarity of stimuli. In our use of INDSCAL, the stimuli are the 75 authors, and the three periods are regarded as three separate ‘‘judges.’’
Usually, of course, persons are the judges in INDSCAL studies, and the ‘‘derived subject weights,’’ which are standard INDSCAL output, are taken to show the salience of each dimension to each person. In replacing individuals as judges with large numbers of citers, we are acting as if the citers collectively embodied the paradigm of information science in each 8-year period.
Accordingly, we interpret the derived subject weights for each period as indicating the relative importance of the dimensions within the paradigm. Thus, we can probe a hidden aspect of disciplinary history—whether key dimensions of the field were given about the same weight in all periods. If not, that would be consistent with a perception of paradigm shift.
Substantively, it is as if during 1972–1979 citers had regarded the range of specialties as by far the most important part of the information science paradigm, but then during 1980–1987 had taken much more cognizance of the differences in authors’ orientation toward literatures or users.

Perhaps the main weakness of this INDSCAL measure is that it is so indirect—that is, not clearly connected to specific papers with specific claims about the world. One expects evidence of paradigm shifts to leap from main texts, not references; from writers, not citers.
Though it might be used to discover paradigm shift, we think it has more promise as a means ofconfirming one. ... A shift detectable there implies not only that authors are promoting new lines of inquiry, but that citers are responding in such a way that the overall map of the discipline is changed.

Toward that account, ACA simultaneously provides both breadth and focus. It provides breadth by forcing contemplation of multiple specialties... It provides focus by forcing contemplation of particular authors, which is to say particular oeuvres and works. It also provides crude but unmistakable evidence of intellectual change.

The role of information science is to explicate the conceptual and methodological foundations on which existing systems are based’’ (Borko, 1968, p. 67). Or ‘‘Information science is the study of the means by which organised structures (which we call ‘information systems’) process recorded symbols to meet their defined objectives’’ (Hayes, 1985, p. 174) .
What they do study empirically, and uniquely, are problems associated with the human–literature barrier—the special difficulties of obtaining answers to questions from publications, in any medium, rather than persons. In other words, while many scholars seek to understand communication between persons, information scientists seek to understand communication between persons and certain valued surrogates for persons that literatures comprise (White, 1992).
This study requires a conceptual scheme that encompasses properties not only of literatures(e.g., size, growth rate, age, dispersion, authority levels, degree of summarization, quality of indexing) but also of people (e.g., interests and concerns, vocabularies, social ties, knowledge of existing systems, search styles, editorial strategies, resource environments).
The bond between domain analysts and retrievalists is their common interest in the literature barrier and related phenomena on both sides. The barrier in action is exemplified by information overload and underload—recurring topics for authors in both subdisciplines because they require both literatures and users to be discussed in a single framework, as implied by the second dimension of our maps.