21st International Conference on Computing in High Energy and Nuclear Physics | |
Disk storage management for LHCb based on Data Popularity estimator | |
物理学;计算机科学 | |
Hushchyn, Mikhail^1 ; Charpentier, Philippe^2 ; Ustyuzhanin, Andrey^3 | |
Yandex School of Data Analysis, Moscow Institute of Physics and Technology, Russia^1 | |
CERN - European Organization for Nuclear Research, Switzerland^2 | |
Yandex School of Data Analysis, Russia National Research University, Higher School of Economics (HSE), Russia NRC kurchatov Institute, Russia Moscow Institute of Physics and Technology, Russia^3 | |
关键词: Data distribution; Data popularity; Data storage systems; Loss functions; Optimal number; Recommendation reports; Regression algorithms; Usage history; | |
Others : https://iopscience.iop.org/article/10.1088/1742-6596/664/4/042026/pdf DOI : 10.1088/1742-6596/664/4/042026 |
|
学科分类:计算机科学(综合) | |
来源: IOP | |
【 摘 要 】
This paper presents an algorithm providing recommendations for optimizing the LHCb data storage. The LHCb data storage system is a hybrid system. All datasets are kept as archives on magnetic tapes. The most popular datasets are kept on disks. The algorithm takes the dataset usage history and metadata (size, type, configuration etc.) to generate a recommendation report. This article presents how we use machine learning algorithms to predict future data popularity. Using these predictions it is possible to estimate which datasets should be removed from disk. We use regression algorithms and time series analysis to find the optimal number of replicas for datasets that are kept on disk. Based on the data popularity and the number of replicas optimization, the algorithm minimizes a loss function to find the optimal data distribution. The loss function represents all requirements for data distribution in the data storage system. We demonstrate how our algorithm helps to save disk space and to reduce waiting times for jobs using this data.
【 预 览 】
Files | Size | Format | View |
---|---|---|---|
Disk storage management for LHCb based on Data Popularity estimator | 842KB | download |