The goal of multi-document text summarizing is to automatically generate a succinct yet thorough summary from
multiple connected sources. The growing volume of digital text data is making the task of document summarization more
challenging. More effective summarization is now possible because to significant advancements in artificial intelligence
techniques and pre-trained language models; nonetheless, the majority of deep learning approaches require a significant amount
of processing power and large training datasets. This work suggests a multi-document text summarizing method powered by
artificial intelligence that makes use of supervised learning and sentence feature modeling.
The proposed system begins with heavy preprocessing that involves lemmatization, normalization, stop-word removal, and
sentence segmentation. Statistical and structural features, including the TF–IDF representation, length of the sentence,
normalized position of the sentence, and numeric token count, are extracted to denote sentence significance. Like other
supervised extraction algorithms discussed in previous literature, a Logistic Regression classifier is created using a balanced
dataset to classify sentences into significant or insignificant categories. Probabilities are used for ranking of selected sentences,
while cosine similarity is applied under a certain word limit constraint to avoidredundancies.
The suggested approach demonstrates effective accuracy while needing very little processing effort, according to experimental
examination of the DUC 2003 corpus. Based on the ROUGE measure, the model produces ROUGE-1, ROUGE-2, and
ROUGE-L scores of 0.1722, 0.0306, and 0.1048, respectively. It demonstrates how basic AI models combined with useful
sentence-based features could be a reliable benchmark for information extraction from several publications. To increase the
effectiveness of summaries, the authors intend to look into contextual representation algorithms like Sentence-BERT and
implement human evaluation standards in subsequent studies.
[1] K. Shinde, T. Roy, and T. Ghosal, “An extractive-abstractive approach for multi-document summarization of scientific
articles for literature review,” in Proc. Workshop on Scholarly Document Processing, 2022.
[2] Y. Liu and M. Lapata, “Text Summarization with Pretrained Encoders,” Proceedings of EMNLP, 2021.
[3] J. Zhang, Y. Zhao, M. Saleh, and P. Liu, “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive
Summarization,” ICML, 2021.
[4] H. Su, J. Jiang, and Y. Huang, “Multi-Granularity Adaptive Extractive Document Summarization with Heterogeneous
Graph Neural Networks,” PeerJ Computer Science, vol. 9, 2023.
[5] N. Moratanch and S. Chitrakala, “A Survey on Extractive Text Summarization,” International Journal of Computer
Applications, 2021.
[6] T. Gao, X. Yao, and D. Chen, “SimCSE: Simple Contrastive Learning of Sentence Embeddings,” EMNLP, 2021.
[7] I. Beltagy, M. Peters, and A. Cohan, “Longformer: The Long-Document Transformer,” ACL, 2021.
[8] A. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” EMNLP, 2021.
[9] C. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” Text Summarization Branches Out, updated
usage widely cited, 2022.
[10] D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., draft, 2023.
[11] K. Papineni et al., “BLEU: A Method for Automatic Evaluation of Machine Translation,” adapted in summarization
evaluation studies, 2022.
[12] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for
Language Understanding,” updated applications in summarization, 2021.
[13] K. R. Varma and V. B. Balamurugan, “Machine learning approaches for extractive text summarization: A
comparative study,” Expert Systems with Applications, vol. 178, 2021.
[14] R. Bandaru and Y. Radhika, “Extractive multi-document text summarization leveraging hybrid semantic
similarity
measures,” International Journal of Advanced Computer Science and Applications, 2022.
[15]Serasiya S, And U. Chauhan, “Abstractive Gujarati Text Summarization Using Sequence-To-Sequence Model and
Attention Mechanism”, Journal of Information Systems Engineering and Management, 2468-4376, 2025, PP: 754- 762
[16] M. Nasari and A. S. Girsang, “Automated multi-document summarization using extractive-abstractive approaches,”
International Journal of Informatics and Communication Technology, 2024.
[17] N. Gu, E. Ash, and R. Hahnloser, “MemSum: Extractive summarization of long documents using multi-step episodic
Markov decision processes,” arXiv, 2021.
[18] M. Chen et al., “SgSum: Transforming multi-document summarization into sub-graph selection,” in Proceedings of
EMNLP, 2021.
[19]Serasiya S, And U. Chauhan. "A Comprehensive Survey on Text Summarization For Indian Languages: Opportunities,
Challenges And Future Prospects." South Eastern European Journal Of Public Health (2025): 1418-1434.
[20] R. Bandaru and Y. Radhika, “Extractive multi-document text summarization leveraging hybrid semantic similarity
measures,” International Journal of Advanced Computer Science and Applications, 2022.