Advanced Data Processing using Machine Learning: A Comparative Study of Predictive Models for Large-Scale Data Analytics


Date Published : 14 September 2026

Contributors

JERALD

JAIN DEEMED-TO- BE UNIVERSITY
Author

PAWAN

Symbiosis International (Deemed University)
Author

Keywords

Predictive Analytics Large-Scale Data Analytics Feature Engineering Support Vector Machine Predictive Modeling

Proceeding

Track

General Track

License

Copyright (c) 2026 Sustainable Global Societies Initiative

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Abstract

As the amount of data collected in various fields continues to grow, the challenge of deriving meaningful information for making sound decisions is growing apace. The technology of advanced data processing, combined with machine learning (ML), has proven to be an effective approach to analytics, pattern recognition and accurate predictions. In this paper, some of the popular ML models are compared in the context of predictive analytics to large volume data environment. The proposed workflow is as follows: Data preparation, feature selection and transformation, model training, and evaluation of the results to determine the effectiveness of the five supervised classification methods that are used in the study (Logistic Regression, Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost)). They are evaluated based on accuracy, precision, recall, F1 score, and training time of the model for model performance, as well as predictive power, and scalability in terms of efficiency. Moreover, the study examines the validity of the models for real life scenarios, such as trend forecasting, recommendation systems, predicting customer greed and decision support systems. The results have practical suggestions for selection of suitable machine learning algorithms based on your output data volume, and based on your application requirements. The comparative study supports the development of intelligent, data-driven systems, by specifying the advantages and disadvantages of the available predictive models, and by pointing towards future development of scalable and efficient machine learning based data analysis.

References

No References

Downloads

How to Cite

S, J. N. K. ., & Pawan Kumar Verma, P. K. V. (2026). Advanced Data Processing using Machine Learning: A Comparative Study of Predictive Models for Large-Scale Data Analytics. Sustainable Global Societies Initiative, 1(11). https://vectmag.com/sgsi/paper/view/1271