A Statistical-ML Hybrid Approach for Robust Selectivity Estimation on Skewed Data
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Nazarbayev University School of Engineering and Digital Sciences
Abstract
Selectivity estimation is a major challenge in modern database management systems
and can be considered as one of the bottlenecks in query optimization. Traditional
histogram-based summaries often lead to estimation errors when applied on skewed
data. Newly appeared machine learning algorithms can improve the accuracy but
their usage leads to the higher inference latency and training costs, which limits their
capability to be useful in high loaded systems.
We propose Hybrid FD CDF framework for high precision, single attribute, range
selectivity calculation and managing the estimator adaptability for new data and
workload. This framework uses the divide-and-conquer strategy, partitioning the data
domain using Freedman Diaconis Rule to calculate optimal number of buckets with
fixed width. Then, using a waterfall model selection policy select a model that takes
the least resources to reach high accuracy, leading to optimal resource management
and optimization.
Experimental results on synthetically generated datasets with hard distribution
and on real benchmarks show that our approach achieves lower Q-error and inference
time compared to the Equi-Width approach.
Description
Keywords
Citation
Unaspekov, T. (2026). A Statistical-ML Hybrid Approach for Robust Selectivity Estimation on Skewed Data. Nazarbayev University School of Engineering and Digital Sciences
Collections
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwised noted, this item's license is described as Attribution 3.0 United States
