Umeå universitets logga

umu.sePublikationer
Ändra sökning
Länk till posten
Permanent länk

Direktlänk
Prakhya, Karthik
Publikationer (3 of 3) Visa alla publikationer
Prakhya, K., Birdal, T. & Yurtsever, A. (2025). Convex formulations for training two-layer ReLU neural networks. In: 13th International Conference on Learning Representations, ICLR 2025: . Paper presented at International Conference on Learning Representations (ICLR), Singapore, April 24-28, 2025 (pp. 30682-30704). Curran Associates, Inc.
Öppna denna publikation i ny flik eller fönster >>Convex formulations for training two-layer ReLU neural networks
2025 (Engelska)Ingår i: 13th International Conference on Learning Representations, ICLR 2025, Curran Associates, Inc., 2025, s. 30682-30704Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Solving non-convex, NP-hard optimization problems is crucial for training machine learning models, including neural networks. However, non-convexity often leads to black-box machine learning models with unclear inner workings. While convex formulations have been used for verifying neural network robustness, their application to training neural networks remains less explored. In response to this challenge, we reformulate the problem of training infinite-width two-layer ReLU networks as a convex completely positive program in a finite-dimensional (lifted) space. Despite the convexity, solving this problem remains NP-hard due to the complete positivity constraint. To overcome this challenge, we introduce a semidefinite relaxation that can be solved in polynomial time. We then experimentally evaluate the tightness of this relaxation, demonstrating its competitive performance in test accuracy across a range of classification tasks.

Ort, förlag, år, upplaga, sidor
Curran Associates, Inc., 2025
Nyckelord
copositive programming, semidefinite programming, neural networks
Nationell ämneskategori
Artificiell intelligens
Identifikatorer
urn:nbn:se:umu:diva-236599 (URN)2-s2.0-105010230817 (Scopus ID)979-8-3313-2085-0 (ISBN)
Konferens
International Conference on Learning Representations (ICLR), Singapore, April 24-28, 2025
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Knut och Alice Wallenbergs StiftelseVetenskapsrådet
Tillgänglig från: 2025-03-17 Skapad: 2025-03-17 Senast uppdaterad: 2025-07-18Bibliografiskt granskad
Dadras, A., Banerjee, S., Prakhya, K. & Yurtsever, A. (2024). Federated Frank-Wolfe algorithm. In: Albert Bifet; Jesse Davis; Tomas Krilavičius; Meelis Kull; Eirini Ntoutsi; Indrė Žliobaitė (Ed.), Machine learning and knowledge discovery in databases. Research track: European Conference, ECML PKDD 2024, Vilnius, Lithuania, September 9–13, 2024, proceedings, part III. Paper presented at European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2024), Vilnius, Lithuania, September 9-13, 2024 (pp. 58-75). Springer Nature
Öppna denna publikation i ny flik eller fönster >>Federated Frank-Wolfe algorithm
2024 (Engelska)Ingår i: Machine learning and knowledge discovery in databases. Research track: European Conference, ECML PKDD 2024, Vilnius, Lithuania, September 9–13, 2024, proceedings, part III / [ed] Albert Bifet; Jesse Davis; Tomas Krilavičius; Meelis Kull; Eirini Ntoutsi; Indrė Žliobaitė, Springer Nature, 2024, s. 58-75Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Federated learning (FL) has gained a lot of attention in recent years for building privacy-preserving collaborative learning systems. However, FL algorithms for constrained machine learning problems are still limited, particularly when the projection step is costly. To this end, we propose a Federated Frank-Wolfe Algorithm (FedFW). FedFW features data privacy, low per-iteration cost, and communication of sparse signals. In the deterministic setting, FedFW achieves an ε-suboptimal solution within O(ε-2) iterations for smooth and convex objectives, and O(ε-3) iterations for smooth but non-convex objectives. Furthermore, we present a stochastic variant of FedFW and show that it finds a solution within O(ε-3) iterations in the convex setting. We demonstrate the empirical performance of FedFW on several machine learning tasks.

Ort, förlag, år, upplaga, sidor
Springer Nature, 2024
Serie
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 14943
Nyckelord
federated learning, frank wolfe, conditional gradient method, projection-free, distributed optimization
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
urn:nbn:se:umu:diva-228614 (URN)10.1007/978-3-031-70352-2_4 (DOI)001308375900004 ()978-3-031-70351-5 (ISBN)978-3-031-70352-2 (ISBN)
Konferens
European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2024), Vilnius, Lithuania, September 9-13, 2024
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Vetenskapsrådet, 2023-05476
Anmärkning

Also part of the book sub series: Lecture Notes in Artificial Intelligence (LNAI). 

Tillgänglig från: 2024-08-19 Skapad: 2024-08-19 Senast uppdaterad: 2025-04-24Bibliografiskt granskad
Dadras, A., Prakhya, K. & Yurtsever, A. (2022). Federated Frank-Wolfe Algorithm. In: : . Paper presented at FL-NeurIPS'22, International Workshop on Federated Learning: Recent Advances and New Challenges in Conjunction with NeurIPS 2022, New Orleans, LA, USA, December 2, 2022.
Öppna denna publikation i ny flik eller fönster >>Federated Frank-Wolfe Algorithm
2022 (Engelska)Konferensbidrag, Poster (med eller utan abstract) (Refereegranskat)
Abstract [en]

Federated learning (FL) has gained much attention in recent years for building privacy-preserving collaborative learning systems. However, FL algorithms for constrained machine learning problems are still very limited, particularly when the projection step is costly. To this end, we propose a Federated Frank-Wolfe Algorithm (FedFW). FedFW provably finds an ε-suboptimal solution of the constrained empirical risk-minimization problem after O(ε−2) iterations if the objective function is convex. The rate becomes O(ε−3) if the objective is non-convex. The method enjoys data privacy, low per-iteration cost and communication of sparse signals. We demonstrate empirical performance of the FedFW algorithm on several machine learning tasks.

Nyckelord
federated learning, frank wolfe, conditional gradient method, projection-free, distributed optimization
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
urn:nbn:se:umu:diva-205126 (URN)
Konferens
FL-NeurIPS'22, International Workshop on Federated Learning: Recent Advances and New Challenges in Conjunction with NeurIPS 2022, New Orleans, LA, USA, December 2, 2022
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Tillgänglig från: 2023-02-23 Skapad: 2023-02-23 Senast uppdaterad: 2025-01-23Bibliografiskt granskad
Organisationer

Sök vidare i DiVA

Visa alla publikationer