CONTEXT-AWARE AUTOMATIC TERM EXTRACTION USING USER BEHAVIOR AND TEXTUAL FEATURES

Authors

  • Vineetha Singa Assistant Professor, Department of Master of Computer Application, Guru Nanak Dev Engineering College, Bidar, India
  • Vigneshwar 2nd Semester, Department of Master of Computer Applications, Guru Nanak Dev Engineering College, Bidar, India
  • Ashwin 2nd Semester, Department of Master of Computer Applications, Guru Nanak Dev Engineering College, Bidar, India

Keywords:

Automatic Term Extraction, Context-Aware NLP, User Behavior Analysis, TF-IDF, Hybrid Scoring, Information Retrieval, Knowledge Discover

Abstract

Automatic Term Extraction (ATE) is a pivotal task in Natural Language Processing (NLP) that seeks to identify semantically significant words and multi-word expressions from domain-specific corpora. Classical ATE methods, grounded in statistical heuristics such as Term Frequency–Inverse Document Frequency (TF-IDF) or purely linguistic pipelines, often fail to reflect the genuine importance of terms in interactive, user-driven environments. This paper proposes a context-aware ATE framework that synergistically integrates textual features—including term frequency, positional salience, syntactic patterns, and co-occurrence graphs—with implicit user behavior signals such as query history, dwell time, click-through rate, and session-level engagement metrics. A hybrid scoring function merges these heterogeneous signals into a unified relevance score for each candidate term. The system is evaluated on a benchmark corpus drawn from the ACM Digital Library and a proprietary e-learning dataset. Experimental results demonstrate that the proposed approach achieves a Precision of 0.87, Recall of 0.83, and F1-Score of 0.85, outperforming baseline TF-IDF and RAKE-based systems by 9–14% in F1. The framework is lightweight, language-agnostic at the feature level, and readily deployable in real-time digital library, search-engine, and recommendation-system pipelines.

References

I. S. Rigouts Terryn, V. Hoste, and E. Lefever, "Looking for Terms in Corpora with HAMLET," Terminology, vol. 28, no. 1, pp. 1–34, 2022.

II. K. T. Frantzi, S. Ananiadou, and H. Mima, "Automatic recognition of multi-word terms: the C-value/NC-value method," International Journal on Digital Libraries, vol. 3, no. 2, pp. 115–130, 2000.

III. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in Proc. NAACL-HLT 2019, pp. 4171–4186.

IV. I. Beltagy, K. Lo, and A. Cohan, "SciBERT: A Pretrained Language Model for Scientific Text," in Proc. EMNLP 2019, pp. 3615–3620.

V. D. Kelly and J. Teevan, "Implicit Feedback for Inferring User Preference: A Bibliography," ACM SIGIR Forum, vol. 37, no. 2, pp. 18–28, 2003.

VI. H. Nakagawa and T. Mori, "A Simple but Powerful Automatic Term Extraction Method," in Proc. COLING-02 Workshop on Computational Terminology, 2002.

VII. Y. Zhou, W. Lam, and D. Lee, "Domain Specific Term Extraction Using Domain-Relevance Measures," in Proc. CIKM 2010, pp. 1429–1432.

VIII. B. Daille, "Terminology Mining," in Mining Text Data, Springer, pp. 291–327, 2012.

IX. M. S. Conrado, T. A. S. Pardo, and S. O. Rezende, "A Machine Learning Approach to Automatic Term Extraction," in Proc. HLT-NAACL 2013 Workshop, pp. 9–17.

X. D. Kelly and J. Teevan, "Implicit Feedback for Inferring User Preference," ACM SIGIR Forum, vol. 37, no. 2, pp. 18–28, 2003.

XI. T. Joachims et al., "Accurately Interpreting Clickthrough Data as Implicit Feedback," in Proc. SIGIR 2005, pp. 154–161.

XII. X. Yi et al., "Beyond Clicks: Dwell Time for Personalization," in Proc. ACM RecSys 2014, pp. 113–120.

XIII. A. Dey, "Understanding and Using Context," Personal and Ubiquitous Computing, vol. 5, no. 1, pp. 4–7, 2001.

XIV. Y. Zhang, L. Liu, and M. Song, "Context-Aware Keyphrase Extraction in E-Learning Environments," Expert Systems with Applications, vol. 215, 119416, 2023.

XV. F. Boudin, "Unsupervised Keyphrase Extraction with Multipartite Graphs," in Proc. NAACL 2018, pp. 667–672.

XVI. M. Honnibal and I. Montani, "spaCy 3: Natural Language Understanding," Explosion AI, 2023.

XVII. C. Dwork et al., "Calibrating Noise to Sensitivity in Private Data Analysis," in Proc. TCC 2006, pp. 265–284.

XVIII. S. Rose et al., "Automatic Keyword Extraction from Individual Documents," in Text Mining: Applications and Theory, Wiley, pp. 1–20, 2010.

XIX. P. Kairouz et al., "Advances and Open Problems in Federated Learning," Foundations and Trends in Machine Learning, vol. 14, no. 1–2, pp. 1–210, 2021.

XX. M. Krapivin, A. Autaeu, and M. Marchese, "Large Dataset for Keyphrases Extraction," University of Trento Technical Report DISI-09-055, 2009.

XXI. R. Navigli and P. Velardi, "Learning Word-Class Lattices for Definition and Hypernym Extraction," in Proc. ACL 2010, pp. 1318–1327.

XXII. S. Arora, Y. Liang, and T. Ma, "A Simple but Tough-to-Beat Baseline for Sentence Embeddings," in Proc. ICLR 2017.

XXIII. L. Page, S. Brin, R. Motwani, and T. Winograd, "The PageRank Citation Ranking: Bringing Order to the Web," Stanford Technical Report 1999-66, 1999.

XXIV. R. Mihalcea and P. Tarau, "TextRank: Bringing Order into Text," in Proc. EMNLP 2004, pp. 404–411.

XXV. A. Siddiqi and A. Sharan, "Keyword and Keyphrase Extraction Techniques: A Literature Review," International Journal of Computer Applications, vol. 109, no. 2, pp. 18–23, 2015.

XXVI. T. Kudo and J. Richardson, "Sentence Piece," in Proc. EMNLP 2018 Demo, pp. 66–71.

XXVII. A. Vaswani et al., "Attention Is All You Need," in Proc. NeurIPS 2017, pp. 5998–6008.

XXVIII. Y. Liu et al., "RoBERTa: A Robustly Optimized BERT Pretraining Approach," arXiv:1907.11692, 2019.

XXIX. M. Krapivin, A. Autaeu, and M. Marchese, "Large Dataset for Keyphrases Extraction," University of Trento Technical Report DISI-09-055, 2009.

XXX. S. M. Kim, O. Valenzuela-Escarcega, and E. H. Hovy, "A Survey of Automatic Term Extraction," ACM Computing Surveys, vol. 55, no. 4, pp. 1–36, 2022.

XXXI. G. Adomavicius and A. Tuzhilin, "Context-Aware Recommender Systems," in Recommender Systems Handbook, Springer, 3rd ed., pp. 217–253, 2022.

XXXII. P. Lample and F. Conneau, "Cross-lingual Language Model Pretraining," in Proc. NeurIPS 2019, pp. 7059–7069.

XXXIII. S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python. O'Reilly Media, 2009.

XXXIV. C. D. Manning, P. Raghavan, and H. Schutze, Introduction to Information Retrieval. Cambridge University Press, 2008.

XXXV. X. Qian, Y. Liu, and F. Ye, "User Behavior-Enhanced Term Weighting for Personalized Information Retrieval," Information Processing and Management, vol. 60, no. 3, 103285, 2023.

Additional Files

Published

01-06-2026

How to Cite

Vineetha Singa, Vigneshwar, & Ashwin. (2026). CONTEXT-AWARE AUTOMATIC TERM EXTRACTION USING USER BEHAVIOR AND TEXTUAL FEATURES. International Educational Journal of Science and Engineering, 9(05), 619–626. Retrieved from https://iejse.com/journals/index.php/iejse/article/view/401