Umeå universitets logga

umu.sePublikationer
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Missing data and the preprocessing perceptron
Umeå universitet, Teknisk-naturvetenskapliga fakulteten, Institutionen för datavetenskap.
2004 (Engelska)Rapport (Övrigt vetenskapligt)
Abstract [en]

In this paper, several ways to handle missing data, e.g. removing cases, mean imputation, and multiple imputation, are described and discussed. The Pima-Indians-Diabetes data set is used as a case study. This particular data set is interesting to use since it has not been obvious to all users that it actually contains a substantial amount of missing data. The data set is described in detail and the methods for coping with missing data mentioned in the text is applied on the data set.

The preprocessing perceptron is used to train decision support systems on the data sets. A sketch of a way to impute missing data using the preprocessing perceptron is also proposed and discussed. The accuracy of the trained decision support systems, at the optimal efficiency point, lied in the interval 76-82% for the different methods. The highest values were obtained when all missing data cases were removed both from the test and the training set. This is, however, not a good way to handle missing data since the resulting decision support system is biased. Furthermore it will not be able to handle missing data when used on real data in the future. The results of the remaining methods were surprisingly similar, a reason for this might be that the data set used is rather large. Differences between methods would probably be larger in a smaller data set with larger amount of missing data.

Ort, förlag, år, upplaga, sidor
Umeå University, 2004. , s. 26
Serie
UMINF ; 04.02
Nationell ämneskategori
Datorsystem
Forskningsämne
administrativ databehandling
Identifikatorer
URN: urn:nbn:se:umu:diva-8399OAI: oai:DiVA.org:umu-8399DiVA, id: diva2:148070
Tillgänglig från: 2008-01-21 Skapad: 2008-01-21 Senast uppdaterad: 2024-07-04Bibliografiskt granskad

Open Access i DiVA

fulltext(794 kB)33 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 794 kBChecksumma SHA-512
ccc9a88cebc7a13f637db535be4b0ae9737c9335eb7438dcfdca7e446a3ad1ed7bdf039d7263ae2d122e40511855bafae787ff7bde2369fe984cdd673a68c3f8
Typ fulltextMimetyp application/pdf

Person

Kallin Westin, Lena

Sök vidare i DiVA

Av författaren/redaktören
Kallin Westin, Lena
Av organisationen
Institutionen för datavetenskap
Datorsystem

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 33 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

urn-nbn

Altmetricpoäng

urn-nbn
Totalt: 503 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf