A Statistical-ML Hybrid Approach for Robust Selectivity Estimation on Skewed Data

Abstract

Selectivity estimation is a major challenge in modern database management systems and can be considered as one of the bottlenecks in query optimization. Traditional histogram-based summaries often lead to estimation errors when applied on skewed data. Newly appeared machine learning algorithms can improve the accuracy but their usage leads to the higher inference latency and training costs, which limits their capability to be useful in high loaded systems. We propose Hybrid FD CDF framework for high precision, single attribute, range selectivity calculation and managing the estimator adaptability for new data and workload. This framework uses the divide-and-conquer strategy, partitioning the data domain using Freedman Diaconis Rule to calculate optimal number of buckets with fixed width. Then, using a waterfall model selection policy select a model that takes the least resources to reach high accuracy, leading to optimal resource management and optimization. Experimental results on synthetically generated datasets with hard distribution and on real benchmarks show that our approach achieves lower Q-error and inference time compared to the Equi-Width approach.

Description

Keywords

Citation

Unaspekov, T. (2026). A Statistical-ML Hybrid Approach for Robust Selectivity Estimation on Skewed Data. Nazarbayev University School of Engineering and Digital Sciences

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Except where otherwised noted, this item's license is described as Attribution 3.0 United States