A Hybrid Ensemble Framework for Petroleum Licensing Decision Support: Comparing Random Forest, Gradient Boosting, and Their Combination

Authors

  • Apollo Otemo Muchilwa Jomo Kenyatta University of Agriculture and Technology
  • Dr. Kennedy Ogada Jomo Kenyatta University of Agriculture and Technology
  • Dr. Jael Wekesa Jomo Kenyatta University of Agriculture and Technology
  • Newton Malongo The Technical University of Kenya

DOI:

https://doi.org/10.47604/ijts.3944

Keywords:

Petroleum Licensing, Predictive Modeling, Random Forest, Gradient Boosting, Hybrid Ensemble, Decision Support Systems

Abstract

Purpose: This study developed a risk-aware petroleum-licensing decision-support framework centered on a transparency-constrained soft-vote protocol that balances predictive performance with regulatory auditability.

Methodology: Random Forest (RF) and Gradient Boosting Machines (GBM) were evaluated in an empirical laboratory proof of concept using a synthetically generated regulatory dataset (N = 2,723), a stratified 70/30 split (n_train = 1,906; n_test = 817), ten-fold cross-validation, and accuracy, precision, recall, F1-score, ROC-AUC, and pairwise McNemar tests. Candidate RF weights from 0 to 1 were assessed, with equal weighting retained when its F1-score was within 0.005 of the optimum.

Findings: RF and GBM each achieved 98.53% test accuracy, while the hybrid achieved 98.41% accuracy, an F1-score of 97.63%, and a ROC-AUC of 0.9988. Pairwise accuracy differences were not statistically significant. The hybrid's main advantage was risk stabilization: it produced the highest mean cross-validation F1-score (0.985) and the lowest fold-to-fold variability (SD = 0.004), while preserving an auditable equal-weight rule.

Unique Contribution to Theory, Practice and Policy: The framework should be validated on real multi-jurisdictional licensing records before deployment, with jurisdiction-specific asymmetric cost matrices, probability calibration, SHAP or LIME explanations, temporal validation, and mandatory human oversight.

Downloads

Download data is not yet available.

References

Ala’a, M., Qasim, A., Hameed, S., & Fattah, J. (2020). Predictive analytics using machine learning for oil and gas data. Energies, 13(16), 4082.

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.

Callens, A., Morichon, D., Abadie, S., Delpey, M., & Liquet, B. (2020). Using random forest and gradient boosting trees to improve wave forecast at a specific location. Applied Ocean Research, 104, 102339.

Cherepovitsyn, A., Rutenko, E., & Solovyova, V. (2021). Sustainable development of oil and gas resources: A system of environmental, socio-economic, and innovation indicators. Journal of Marine Science and Engineering, 9(11), 1307.

International Energy Agency. (2021). Digitalization and energy. https://www.iea.org/reports/digitalization-and-energy

James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning with applications in R (2nd ed.). Springer.

Josso, P., Bertrand, G., Marcoux, E., & Cassard, D. (2023). Application of random forest for mineral prospectivity mapping: A case study of ferromanganese crust deposits. Ore Geology Reviews, 155, 105347.

Kuhn, M., & Johnson, K. (2020). Feature engineering and selection: A practical approach for predictive models. CRC Press.

Kumar, R., Singh, P., & Verma, A. (2023). Decision tree-based predictive modeling for petroleum licensing outcomes. Journal of Energy Policy Research, 10(2), 145–159.

Mienye, I. D., & Sun, Y. (2022). A survey of ensemble learning: Concepts, algorithms, applications, and prospects. IEEE Access, 10, 99129–99149.

Petroleum Economist. (2022). The state of predictive analytics in the oil and gas industry. https://www.petroleum-economist.com/articles/markets/trends/2022/the-state-of-predictive-analytics-in-the-oil-and-gas-industry

Platt, J. C. (1999). Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In A. J. Smola, P. Bartlett, B. Schölkopf, & D. Schuurmans (Eds.), Advances in large margin classifiers (pp. 61–74). MIT Press.

Probst, P., Wright, M. N., & Boulesteix, A. L. (2020). Hyperparameters and tuning strategies for random forest. WIREs Data Mining and Knowledge Discovery, 9(3), e1301.

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144. https://doi.org/10.1145/2939672.2939778

Sircar, A., Yadav, K., Rayavarapu, K., Bist, N., & Oza, H. (2021). Application of machine learning and artificial intelligence in the oil and gas industry. Petroleum Research, 6(4), 379–391.

Soltani, F., Hajian, M., Toghraie, D., Gheisari, A., Sina, N., & Alizadeh, A. A. (2021). Applying Artificial Neural Networks (ANNs) for prediction of the thermal characteristics of engine oil–based nanofluids containing tungsten oxide-MWCNTs. Case Studies in Thermal Engineering, 26, 101122.

Tuboalabo, A., Buinwi, J. A., Buinwi, U., Okatta, C. G., & Johnson, E. (2024). Leveraging business analytics for competitive advantage: Predictive models and data-driven decision making. International Journal of Management & Entrepreneurship Research, 6(6), 1997–2014.

Wang, Z., Cheng, Z., Ding, X., & Xia, L. (2024). Research on intelligent decision support systems for oil and gas exploration based on machine learning. Plos one, 19(12), e0314108.

Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259. https://doi.org/10.1016/S0893-6080(05)80023-1

Wu, C., Wang, S., Yuan, J., Li, C., & Zhang, Q. (2020). A prediction model of specific productivity index using least square support vector machine method. Advances in Geo-Energy Research, 4(4), 460-467.

Zhang, H., Zimmerman, J., Nettleton, D., & Nordman, D. J. (2020). Random forest prediction intervals. The American Statistician.

Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 694–699. https://doi.org/10.1145/775047.775151

Downloads

Published

2026-08-25

How to Cite

Muchilwa, A., Ogada, K., Wekesa, J., & Malongo, N. (2026). A Hybrid Ensemble Framework for Petroleum Licensing Decision Support: Comparing Random Forest, Gradient Boosting, and Their Combination. International Journal of Technology and Systems, 11(1), 83–116. https://doi.org/10.47604/ijts.3944

Issue

Section

Articles