An Empirical Study on the Impact of Feature Scaling and Encoding Strategies on Machine Learning Regression Pipelines
DOI:
https://doi.org/10.47738/ijiis.v9i1.293Keywords:
Machine Learning Regression, Data Preprocessing, Feature Scaling, Categorical Encoding, Pipeline EvaluationAbstract
Data preprocessing is a critical yet often underestimated component of Machine Learning (ML) regression pipelines. While prior studies have largely focused on algorithm selection and model architecture, the combined impact of feature scaling and categorical encoding strategies within end-to-end regression pipelines remains insufficiently explored. This study presents an empirical evaluation of how different preprocessing configurations influence regression model performance. Three regression algorithms, Linear Regression, Random Forest Regression, and Gradient Boosting Regression are evaluated in combination with multiple feature scaling methods (Min–Max, Standard, and Robust scaling) and categorical encoding techniques (One-Hot and Ordinal encoding). Experiments are conducted on a real-world car sales dataset comprising 50,000 records, using a k-fold cross-validation framework to ensure robust performance estimation. Model performance is assessed primarily using mean R², supported by RMSE and MAE as error-based metrics. The results demonstrate that ensemble-based models, particularly Gradient Boosting and Random Forest, consistently outperform Linear Regression across all preprocessing configurations. Feature scaling shows limited influence on ensemble model performance, whereas categorical encoding plays a more significant role, with One-Hot Encoding yielding higher predictive accuracy and lower error dispersion than Ordinal Encoding. Overall, the findings highlight that model choice is the dominant determinant of regression performance, followed by encoding strategy, while scaling has a comparatively minor effect. This study provides empirical guidance for designing robust and effective ML regression pipelines and underscores the importance of evaluating preprocessing techniques in conjunction with model selection.References
G. Erol, B. Uzbaş, C. Yücelbaş, and Ş. Yücelbaş, “Analyzing the Effect of Data Preprocessing Techniques Using Machine Learning Algorithms on the Diagnosis of COVID‐19,” Concurr. Comput. Pract. Exp., vol. 34, no. 28, pp. e7393, 2022, doi: 10.1002/cpe.7393.
O. G. Emi-Johnson and K. J. Nkrumah, “Predicting 30-Day Hospital Readmission in Patients with Diabetes Using Machine Learning on Electronic Health Record Data,” Cureus, vol. 17, no. 4, pp. e82437, 2025, doi: 10.7759/cureus.82437.
A. Alagic, N. Zivic, E. Kadusic, D. Hamzic, N. Hadzajlic, M. Dizdarevic, and E. Selmanovic, “Machine Learning for an Enhanced Credit Risk Analysis: A Comparative Study of Loan Approval Prediction Models Integrating Mental Health Data,” Mach. Learn. Knowl. Extr., vol. 6, no. 1, pp. 53–77, 2024, doi: 10.3390/make6010004.
K. Parveen, “Enhanced Credit Scoring Prediction Using KNN-Z-Score Based Logistic Regression (KZ-LR) Algorithm,” J. Electr. Syst., vol. 20, no. 3, pp. 7230–7237, 2024, doi: 10.52783/jes.7419.
P. Yasodha, “Data Preprocessing Methods for Machine Learning: An Empirical Comparison,” Int. J. Multidiscip. Res., vol. 7, no. 3, pp. 1–7, 2025, doi: 10.36948/ijfmr.2025.v07i03.48569.
S. Ramya, B. Kumaraswamy, V. Agarwal, and A. K. Jain, “A Comparative Study of Different Data Pre-Processing Methods for Machine Learning,” Int. J. Multidiscip. Res., vol. 7, no. 4, pp. 1–8, 2025, doi: 10.36948/ijfmr.2025.v07i04.52920.
H. Ouyang, L. Tang, J. Ma, and T. Pang, “Application of Hyperspectral Technology with Machine Learning for Brix Detection of Pastry Pears,” Plants, vol. 13, no. 8, pp. 1163, 2024, doi: 10.3390/plants13081163.
I. Fayyaz, G. G. M. N. Ali, and S. S. Khairunnesa, “Advanced Feature Engineering and Machine Learning Techniques for High Accurate Price Prediction of Heterogeneous Pre-Own Cars,” Vehicles, vol. 7, no. 3, pp. 94, 2025, doi: 10.3390/vehicles7030094.
B. Huang, L. Xue, J. Yang, C. Yin, K. Liao, M. Liu, J. Li, and M. Shang, “Regression Study on Fruit‐setting Days of Purple Eggplant Fruit Based on in Situ VIS‐NIRS and Attention Cycle Neural Network,” J. Food Sci., vol. 90, no. 1, pp. e17593, 2025, doi: 10.1111/1750-3841.17593.
V. Inturi, S. V Balaji, P. Gyanam, B. P. V Pragada, G. R. Sabareesh, and V. Pakrashi, “An Integrated Condition Monitoring Scheme for Health State Identification of a Multi-Stage Gearbox Through Hurst Exponent Estimates,” Struct. Heal. Monit., vol. 22, no. 1, pp. 730–745, 2022, doi: 10.1177/14759217221092828.
E. Tabane, E. Mnkandla, and Z. Wang, “Optimizing DNA Sequence Classification via a Deep Learning Hybrid of LSTM and CNN Architecture,” Appl. Sci., vol. 15, no. 15, pp. 8225, 2025, doi: 10.3390/app15158225.
M. D. Hossain, S. H. Scott, T. Cluff, and S. P. Dukelow, “The Use of Machine Learning and Deep Learning Techniques to Assess Proprioceptive Impairments of the Upper Limb After Stroke,” J. Neuroeng. Rehabil., vol. 20, no. 1, pp. 1–18, 2023, doi: 10.1186/s12984-023-01140-9.
R. Van, D. Alvarez, T. Mize, S. Gannavarapu, L. C. Reddy, F. Nasoz, and M. V. Han, “A Comparison of RNA-Seq Data Preprocessing Pipelines for Transcriptomic Predictions Across Independent Studies,” BMC Bioinformatics, vol. 25, no. 1, pp. 1–22, 2024, doi: 10.1186/s12859-024-05801-x.
Y. Sun, N. S. Nayani, Y. Xu, Z. Xu, J. Yang, and Y. Feng, “Rapid and Nondestructive Determination of Oil Content and Distribution of Potato Chips Using Hyperspectral Imaging and Chemometrics,” ACS Food Sci. Technol., vol. 4, no. 6, pp. 1579–1588, 2024, doi: 10.1021/acsfoodscitech.4c00196.
J. Ong, W. He, P. Maglanque, X. Jiang, L. M. Gillman, A. Vergis, and K. Hardy, “A Preprocessing Pipeline for Pupillometry Signal from Multimodal iMotion Data,” Sensors, vol. 25, no. 15, pp. 4737, 2025, doi: 10.3390/s25154737.
Q. Xiao, J. Zheng, J. Wen, F. Deng, R. Gu, L. Li, Y. He, and J. Yang, “Rapid Detection of Physicochemical Indicators of Tobacco Flavorings Using Fourier-Transform Near Infrared Spectroscopy with Chemometrics and Machine Learning,” ACS Omega, vol. 10, no. 19, pp. 19714–19722, 2025, doi: 10.1021/acsomega.5c00225.
L. J. Keevers and P. Jean-Richard-dit-Bressel, “Obtaining Artifact-Corrected Signals in Fiber Photometry via Isosbestic Signals, Robust Regression, and Calculations,” Neurophotonics, vol. 12, no. 02, pp. 025003-1-025003–12, 2025, doi: 10.1117/1.nph.12.2.025003.
F. Koopmans, K. W. Li, R. V Klaassen, and A. B. Smit, “MS-DAP Platform for Downstream Data Analysis of Label-Free Proteomics Uncovers Optimal Workflows in Benchmark Data Sets and Increased Sensitivity in Analysis of Alzheimer’s Biomarker Data,” J. Proteome Res., vol. 22, no. 2, pp. 374–386, 2022, doi: 10.1021/acs.jproteome.2c00513.
A. Mansoori, M. Zeinalnezhad, and L. Nazarimanesh, “Optimization of Tree-Based Machine Learning Models to Predict the Length of Hospital Stay Using Genetic Algorithm,” J. Healthc. Eng., vol. 2023, no. 1, pp. 1–14, 2023, doi: 10.1155/2023/9673395.
B. Avanzi, G. Taylor, M. Wang, and B. Wong, “Machine Learning with High-Cardinality Categorical Features in Actuarial Applications,” Astin Bull., vol. 54, no. 2, pp. 213–238, 2024, doi: 10.1017/asb.2024.7.
F. Pargent, F. Pfisterer, J. Thomas, and B. Bischl, “Regularized Target Encoding Outperforms Traditional Methods in Supervised Machine Learning with High Cardinality Features,” Comput. Stat., vol. 37, no. 5, pp. 2671–2692, 2022, doi: 10.1007/s00180-022-01207-6.
S. S. L. Parvathi, A. D. B, G. L. Kulkarni, S. Murugan, B. K. P. Vijayammal, and Neha, “Exploring Feature Relationships in Brain Stroke Data Using Polynomial Feature Transformation and Linear Regression Modeling,” J. Mach. Comput., vol. 4, no. 4, pp. 1158–1169, 2024, doi: 10.53759/7669/jmc202404107.
P. Cerda and G. Varoquaux, “Encoding High-Cardinality String Categorical Variables,” IEEE Trans. Knowl. Data Eng., vol. 34, no. 3, pp. 1164–1176, 2022, doi: 10.1109/tkde.2020.2992529.
S. Moseen and B. R. Kumar, “Hybrid Regression Approach for Prediction of Renewable Energy from Human Footstep Power Systems,” Int. J. Data Sci. IoT Manag. Syst., vol. 4, no. 3, pp. 328–340, 2025, doi: 10.64751/ijdim.2025.v4.n3.pp328-340.
R. A. M. Aljohani, “Exploring Football Player Salary Prediction Using Random Forest: Leveraging Player Demographics and Team Associations,” Int. J. Appl. Inf. Manag., vol. 5, no. 4, pp. 203–213, 2025, doi: 10.47738/ijaim.v5i4.115.
D. Zhang, “Task-Based Teaching Method in English Teaching from the Perspective of Big Data,” Int. J. Web-Based Learn. Teach. Technol., vol. 20, no. 1, pp. 1–17, 2025, doi: 10.4018/ijwltt.381309.
W. Jiang, “Key Selection Factors Influencing Animation Films from the Perspective of the Audience,” Mathematics, vol. 12, no. 10, pp. 1–21, 2024, doi: 10.3390/math12101547.
A. L. Aranha, L. L. B. Bernucci, and K. Vasconcelos, “Effects of Different Training Datasets on Machine Learning Models for Pavement Performance Prediction,” Transp. Res. Rec. J. Transp. Res. Board, vol. 2677, no. 8, pp. 196–206, 2023, doi: 10.1177/03611981231155902.
L. Meena and T. Velmurugan, “Optimizing Facial Expression Recognition Through Effective Preprocessing Techniques,” J. Comput. Commun., vol. 11, no. 12, pp. 86–101, 2023, doi: 10.4236/jcc.2023.1112006.
S. S. Samaan and H. A. Jeiad, “Architecting a Machine Learning Pipeline for Online Traffic Classification in Software Defined Networking Using Spark,” IAES Int. J. Artif. Intell., vol. 12, no. 2, pp. 861, 2023, doi: 10.11591/ijai.v12.i2.pp861-873.
N. Maher and S. A. Yousif, “An Automated Machine Learning Model for Diagnosing Coronavirus Disease 2019 (COVID-19) Infection,” IAES Int. J. Artif. Intell., vol. 12, no. 3, pp. 1360, 2023, doi: 10.11591/ijai.v12.i3.pp1360-1369.
P. E. Guillem, M. Zurdo-Tabernero, N. E. Iglesias, Á. Canal-Alonso, L. Durón Figueroa, G. Hernández, A. González-Arrieta, and F. de la Prieta, “Leveraging Transformers for Semi-Supervised Pathogenicity Prediction with Soft Labels,” J. Integr. Bioinform., vol. 22, no. 2, pp. 20240047, 2025, doi: 10.1515/jib-2024-0047.
S. Grafberger, S. Guha, P. Groth, and S. Schelter, “Mlwhatif: What if You Could Stop Re-Implementing Your Machine Learning Pipeline Analyses Over and Over?,” Proc. VLDB Endow., vol. 16, no. 12, pp. 4002–4005, 2023, doi: 10.14778/3611540.3611606.
Downloads
Published
Issue
Section
License
Authors who publish with IJIIS : International Journal on Informatics and Information Systems agree to the following terms: Authors retain copyright and grant the IJIIS : International Journal on Informatics and Information Systems right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share (copy and redistribute the material in any medium or format) and adapt (remix, transform, and build upon the material) the work for any purpose, even commercially with an acknowledgement of the work's authorship and initial publication in IJIIS : International Journal on Informatics and Information Systems. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in IJIIS : International Journal on Informatics and Information Systems. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).

