مقایسه عملکرد رویکردهای کشف و استخراج موضوعات کتاب های الکترونیکی (مقاله علمی وزارت علوم)

درجه علمی: نشریه علمی (وزارت علوم)

نویسندگان: فاطمه زرمهر علی منصوری حسین کارشناس نجف آبادی

منبع: پژوهشنامه پردازش و مدیریت اطلاعات دوره 38 تابستان 1402 شماره 4 (پیاپی 114)

کلید واژه ها: استخراج کلیدواژه های موضوعی متن کاوی مدلسازی موضوعی تجزیه ماتریس نامنفی ماشین بردار پشتیبان کتاب الکترونیکی

حوزه های تخصصی:

حوزه‌های تخصصی علم اطلاعات و دانش‌شناسی

doi: 10.22034/jipm.2023.698598

شماره صفحات: ۱۳۶۹ - ۱۳۹۳

دریافت مقاله تعداد دانلود : ۱۰۰

آرشیو

چکیده

استخراج کلمات کلیدی از مسائل مهم در زمینه پردازش و تحلیل متن بوده و خلاصه ای سطح بالا و دقیق از متن ارائه می دهد. بنابراین انتخاب روش مناسب برای استخراج کلمات کلیدی متن حائز اهمیت است. هدف پژوهش حاضر، مقایسه عملکرد سه رویکرد درکشف و استخراج کلیدواژه های موضوعی کتاب های الکترونیکی با استفاده از تکنیک های متن کاوی و یادگیری ماشین است. در این راستا سه رویکرد آزمایشی شامل: 1.اجرای متوالی فرآیند خوشه بندی، ارتقا کیفیت خوشه ها از نظر معنایی و غنی سازی کلمات توقف حوزه خاص؛ 2. استفاده از الگوی کلیدواژه های تخصصی؛ 3. استفاده از بخش های مهم متن در کشف و استخراج واژگان کلیدی و موضوعات مهم متن، معرفی و مورد مقایسه قرار گرفته است. جامعه آماری، شامل 1000 عنوان کتاب الکترونیکی از زیرشاخه های موضوعی حوزه علم اطلاعات و دانش شناسی بر اساس نظام رده بندی کنگره است که بعد از کسب اطلاعات کتابشناختی آن از پایگاه کتابخانه کنگره، اقدام به تهیه متن اصلی گردید. استخراج کلیدواژهای موضوعی و خوشه بندی داده های آموزش به کمک الگوریتم تجزیه نامنفی ماتریس و با سه رویکرد آزمایشی انجام شد و کیفیت و عملکرد خوشه های موضوعی حاصل از اجرای سه رویکرد در بخش دسته بندی خودکار داده های آزمایشی به کمک ماشین بردار پشتیبان مورد مقایسه قرار گرفت. یافته ها نشان داد افت همینگ (0.020) یا میزان خطا در دسته بندی صحیح متون آزمایشی در رویکرد سوم یعنی بهره گیری از بخش های مهم متن در استخراج کلیدواژه های موضوعی، از دو رویکرد دیگر کمتر است. همچنین امتیاز F1 (0.82) که میانگین دو معیار دقت (0.87) و بازخوانی (0.78) و بازتابی از عملکرد درست فرآیند دسته بندی در برچسب گذاری موضوعی متون است، در رویکرد سوم بهتر از نتایج دو رویکرد دیگر است. نتایج تحلیل ها نشان داد که کیفیت و انسجام معنایی خوشه های موضوعی حاصل از رویکرد سوم یعنی استفاده از بخش های مهم متن در کشف و استخراج موضوع، در مقایسه با دو رویکرد دیگر بهتر بود. بعلاوه کلیدواژه های به دست آمده از خوشه های موضوعی رویکرد سوم را می توان در مجموعه های توصیف نشده و ناشناخته به منظور استخراج محتوای موضوعی ناآشکار کل مجموعه به کار برد.

Comparison of the performance of approaches in discovering and extracting e-book topics

Keyword extraction is one of the most important issues in text processing and analysis and provides a high-level and accurate summary of the text. Therefore, choosing the right method to extract keywords from the text is important. The aim of the present study was to compare the performance of three approaches in discovering and extracting the subject keywords of e-books using text mining and machine learning techniques. In this regard, three experimental approaches have been introduced and compared; including the successive implementation of the clustering process, improving the quality of clusters in terms of semantics and enriching the stop words of a specific field; Use of specialized keyword template; Finally, the use of important parts of the text in discovering and extracting key words and important topics of the text. The statistical population includes 1000 e-book titles from the subject fields of library and information science based on the congress classification system. bibliographic information of EBooks was obtained from the congress library database, then the original text was prepared. The extraction of topic keywords and clustering of training data was performed using the non-negative matrix factorization algorithm with three experimental approaches. The quality and performance of the subject clusters resulting from the implementation of three approaches in the automatic classification of experimental data were compared using a support vector machine. The findings showed that the Hamming loss (0.020) and in other words the error rate in the correct classification of experimental texts in the third approach is far less than the other two approaches. Also, the F1 score (0.82), which is the average of the two criteria of Precision (0.87) and recall (0.78) and is a reflection of the correct performance of the classification process in topic labeling of texts, is better in the third approach than the other two approaches. The results showed that the quality and semantic coherence of the subject clusters obtained from the third approach, ie the use of important parts of the text in discovering and extracting the subject, was better compared to the other two approaches. In this approach, by focusing on the main parts of the data, which represent the main content and theme of the text, more meaningful topic clusters were obtained. In addition, the keywords obtained from the topic cluster of the third approach can be used in unspecified and unknown collections in order to extract the unknown thematic content of the whole collection. The results of third approach also was better in terms of accuracy and readability (0.79) and the rate of classification error (0.020) of texts, in comparison of other two approaches.