Adaptive Neural Pruning for Business-Critical Runtime Optimization

Authors

  • Srilekha Vuyyuru Software Engineer, USA Author

DOI:

https://doi.org/10.15662/IJEETR.2023.0503008

Keywords:

Adaptive Neural Pruning, Deep Neural Networks, Static Pruning

Abstract

The adoption of the Enterprise ML system increases the number of processes that are subject to rigorous constraints, particularly in business-critical applications. These challenges include fraud detection, recommendation engines, access controls, and real-time analytics. While a deep neural network can be achieving high predictive accuracy, the computations are complex, hence resulting in increaseing latency, higher infrastructure costs, and reduced scalability. The restrictions described are particularly related to the environment's impact on real-time responsiveness and cost efficiency, both of which are critical in business. Furthermore, classical compression and pruning are often applied offline and remain static after deployment, limiting their effectiveness in dynamic corporate settings

This study describes an adaptive neural pruning approach for business essential runtime optimizations. The system is  dynamically pruning redundant or low-impact neural components at runtime, based on workload conditions, performance requirements, and even business priorities. Compared to the static pruning strategy, adaptive pruning provides a balance of inference delay, resource consumption, and decision accuracy. The business's important feature preserves a path that enables optimization without compromising high-impact outcomes

The simulation-based experiment will use enterprise-inspired workloads to ensure that latency varies with demand. The findings will  be demonstrating that adaptive neutral pruning is effective in reducing inference time and decision resource use while preserving the accuracy of important decisions. The visualizations, including latency-reduction curves and an accuracy-retention bar chart, demonstrate the efficacy of the approaches. The findings will demonstrate adaptive pruning, a viable and scalable approach that aligns neural network performance with enterprise runtime and cost restrictions

References

1. Yadwadkar, N. J., Romero, F., Li, Q., & Kozyrakis, C. (2019, May). A case for managed and model-less inference serving. In Proceedings of the Workshop on Hot Topics in Operating Systems (pp. 184-191). https://dl.acm.org/doi/abs/10.1145/3317550.3321443

2. Frankle, J., & Carbin, M. (2018). The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635. https://arxiv.org/abs/1803.03635

3. Lemaire, C., Achkar, A., & Jodoin, P. M. (2019). Structured pruning of neural networks with budget-aware regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9108-9116). http://openaccess.thecvf.com/content_CVPR_2019/html/Lemaire_Structured_Pruning_of_Neural_Networks_With_Budget-Aware_Regularization_CVPR_2019_paper.html

4. Tang, S., Li, D., Niu, B., Peng, J., & Zhu, Z. (2019). Sel-INT: A runtime-programmable selective in-band network telemetry system. IEEE transactions on network and service management, 17(2), 708-721. https://ieeexplore.ieee.org/abstract/document/8897503/

5. Lum, K., & Johndrow, J. (2016). A statistical framework for fair predictive algorithms. arXiv preprint arXiv:1610.08077. https://arxiv.org/abs/1610.08077

6. Zhang, Y., Wu, W., Banerjee, S., Kang, J. M., & Sanchez, M. A. (2017, May). SLA-verifier: Stateful and quantitative verification for service chaining. In IEEE INFOCOM 2017-IEEE Conference on Computer Communications (pp. 1-9). IEEE. https://ieeexplore.ieee.org/abstract/document/8057041/

7. Chen, J., & Ran, X. (2019). Deep learning with edge computing: A review. Proceedings of the IEEE, 107(8), 1655-1674. https://ieeexplore.ieee.org/abstract/document/8763885/

Downloads

Published

2023-06-27

How to Cite

Adaptive Neural Pruning for Business-Critical Runtime Optimization. (2023). International Journal of Engineering & Extended Technologies Research (IJEETR), 5(3), 6605-6609. https://doi.org/10.15662/IJEETR.2023.0503008