Umeå University's logo

umu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Subquadratic overparameterization for shallow neural networks
Laboratory for Information and Inference Systems (LIONS), EPFL, Switzerland.
Laboratory for Information and Inference Systems (LIONS), EPFL, Switzerland.
Laboratory for Information and Inference Systems (LIONS), EPFL, Switzerland.
Umeå University, Faculty of Science and Technology, Department of Mathematics and Mathematical Statistics.
Show others and affiliations
2021 (English)In: Advances in neural information processing systems (NeurIPS 2021) / [ed] M. Ranzato; A. Beygelzimer; Y. Dauphin; P.S. Liang; J. Wortman Vaughan, Curran Associates, Inc., 2021, Vol. 34, p. 11247-11259Conference paper, Published paper (Refereed)
Abstract [en]

Overparameterization refers to the important phenomenon where the width of a neural network is chosen such that learning algorithms can provably attain zero loss in nonconvex training. The existing theory establishes such global convergence using various initialization strategies, training modifications, and width scalings. In particular, the state-of-the-art results require the width to scale quadratically with the number of training data under standard initialization strategies used in practice for best generalization performance. In contrast, the most recent results obtain linear scaling either with requiring initializations that lead to the “lazy-training”, or training only a single layer. In this work, we provide an analytical framework that allows us to adopt standard initialization strategies, possibly avoid lazy training, and train all layers simultaneously in basic shallow neural networks while attaining a desirable subquadratic scaling on the network width. We achieve the desiderata via Polyak-Łojasiewicz condition, smoothness, and standard assumptions on data, and use tools from random matrix theory.

Place, publisher, year, edition, pages
Curran Associates, Inc., 2021. Vol. 34, p. 11247-11259
Series
Conference on Neural Information Processing Systems, ISSN 1049-5258
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:umu:diva-196833Scopus ID: 2-s2.0-85131797261ISBN: 9781713845393 (electronic)OAI: oai:DiVA.org:umu-196833DiVA, id: diva2:1673817
Conference
35th Conference on Neural Information Processing Systems, NeurIPS 2021, Virtual (Online), 6 December 6-14, 2021.
Funder
EU, Horizon 2020, 725594EU, European Research CouncilAvailable from: 2022-06-21 Created: 2022-06-21 Last updated: 2022-06-21Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

ScopusPublisher's full text

Authority records

Eftekhari, Armin

Search in DiVA

By author/editor
Eftekhari, Armin
By organisation
Department of Mathematics and Mathematical Statistics
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

isbn
urn-nbn

Altmetric score

isbn
urn-nbn
Total: 245 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf