Umeå universitets logga

umu.sePublikationer
Ändra sökning
Länk till posten
Permanent länk

Direktlänk
Publikationer (10 of 10) Visa alla publikationer
Wallmark, J., Wiberg, M. & Eriksson, M. (2026). Developing and validating a frailty score based on patient-reported outcome 3 months after stroke: a Riksstroke-based study. PLOS ONE, 21(2), Article ID e0343249.
Öppna denna publikation i ny flik eller fönster >>Developing and validating a frailty score based on patient-reported outcome 3 months after stroke: a Riksstroke-based study
2026 (Engelska)Ingår i: PLOS ONE, E-ISSN 1932-6203, Vol. 21, nr 2, artikel-id e0343249Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

BACKGROUND: Frailty is common after stroke and linked to poor outcomes, but many measures are clinician-rated, time-consuming, and not suited to patient-reported data. To address these issues, we developed and validated a frailty score from the Swedish Stroke Register (Riksstroke) three-month follow-up questionnaire.

METHODS: We analyzed responses from 19,470 stroke survivors to nine patient-reported items covering function, mood, fatigue, pain and general health, in the 2021-2022 Riksstroke questionnaire. Dimensionality was assessed with Mokken Scale Analysis and exploratory factor analysis. Item response theory (IRT) was used for score computation. Competing graded response IRT models (unidimensional, correlated-factor, bifactor) were compared, and measurement fairness was examined using differential item functioning (DIF) across age, sex, and education. Prognostic validity was tested with Kaplan-Meier curves and Cox regression for all-cause mortality.

RESULTS: From the Mokken Scale Analysis, all items met scalability criteria. Factor analysis suggested two correlated interpretable facets (Physical Functioning; Well-being/Mental Health). A bifactor IRT model provided the best fit to the data, comprising a general frailty dimension while addressing the strong correlation between the facets. DIF was minimal for sex and education, with modest age-related effects. Higher frailty scores were associated with increased mortality in adjusted Cox models and Kaplan-Meier curves. Tools for computing frailty scores are available at https://github.com/joakimwallmark/frailty-irt-scores.

CONCLUSIONS: A robust, fair, and prognostically meaningful frailty score can be derived from patient-reported items in Riksstroke. More broadly, the study demonstrates how routinely collected patient-reported outcome measures can be leveraged to build scalable frailty scores, offering efficient cost-effective tools for monitoring outcome and guiding quality improvement in stroke care.

Ort, förlag, år, upplaga, sidor
Public Library of Science (PLoS), 2026
Nationell ämneskategori
Medicinsk biostatistik
Forskningsämne
statistik; epidemiologi
Identifikatorer
urn:nbn:se:umu:diva-250646 (URN)10.1371/journal.pone.0343249 (DOI)001695758500001 ()41719319 (PubMedID)2-s2.0-105030698386 (Scopus ID)
Forskningsfinansiär
Vetenskapsrådet, 2022-02046Vetenskapsrådet, 2024-02846
Tillgänglig från: 2026-03-04 Skapad: 2026-03-04 Senast uppdaterad: 2026-03-13Bibliografiskt granskad
Wallmark, J. & Wiberg, M. (2026). The bit scale: a metric score scale for unidimensional item response theory models. Psychometrika, 91(1), 30-46
Öppna denna publikation i ny flik eller fönster >>The bit scale: a metric score scale for unidimensional item response theory models
2026 (Engelska)Ingår i: Psychometrika, ISSN 0033-3123, E-ISSN 1860-0980, Vol. 91, nr 1, s. 30-46Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

In Item Response Theory (IRT), the conventional latent trait scale (θ) is inherently arbitrary, lacking a fixed unit or origin and often tied to specific population distributional assumptions (e.g., standard normal). This limits the direct comparability and interpretability of scores across different tests, populations, or model estimation methods. This paper introduces the “bit scale,” a novel metric transformation for unidimensional IRT scores derived from fundamental principles of information theory, specifically surprisal and entropy. Bit scores are anchored to the properties of the test items rather than the test-taker population. This item-based anchoring ensures the scale’s invariance to population assumptions and provides a consistent metric for comparing latent trait levels. We illustrate the utility of the bit scale through empirical examples: demonstrating consistent scoring when fitting models with different θ scale assumptions, and using anchor items to directly link scores from different test administrations. A simulation study confirms the desirable statistical properties (low bias, accurate standard errors) of Maximum Likelihood estimated bit scores and their robustness to extreme scores. The bit scale offers a theoretically grounded, interpretable, and comparable metric for reporting and analyzing IRT-based assessment results. Software implementations in R (bitscale) and Python (IRTorch) are available and practical implications are discussed.

Ort, förlag, år, upplaga, sidor
Cambridge University Press, 2026
Nyckelord
bit scale, bit scores, information, rescaling, scale transformation
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-248215 (URN)10.1017/psy.2025.10071 (DOI)001658889300001 ()41367367 (PubMedID)2-s2.0-105025666005 (Scopus ID)
Forskningsfinansiär
Marianne och Marcus Wallenbergs Stiftelse, 2019-0129Vetenskapsrådet, 2022-02046
Tillgänglig från: 2026-01-09 Skapad: 2026-01-09 Senast uppdaterad: 2026-05-21Bibliografiskt granskad
Ramsay, J. O., Li, J., Bernstein, C. N., Marrie, R. A., Wallmark, J. & Wiberg, M. (2026). Using the information metric to analyze clinical rating scales. Journal of educational and behavioral statistics, 51(2), 395-418
Öppna denna publikation i ny flik eller fönster >>Using the information metric to analyze clinical rating scales
Visa övriga...
2026 (Engelska)Ingår i: Journal of educational and behavioral statistics, ISSN 1076-9986, E-ISSN 1935-1054, Vol. 51, nr 2, s. 395-418Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

A rating scale is a set of categories designed to obtain information about a quantitative or a qualitative attribute. Item response theory (IRT) proposes that a probability function over a single latent variable represents the overall attribute evolution that the scale is designed to assess. Here we utilize an information theory approach to IRT to analyze rating scale data. The proposed IRT analyses, based on surprisal, offer new tools for assessing raters, rated items, and the whole rating scale. The information transformation from probability to surprisal is a new lens from which to view choice data and is an important augmentation of probability-based IRT. It also offers new graphical tools to measure the amount of information captured by an item in an additive metric, and to measure covariation among items using mutual information. The proposed methodology is illustrated using two scales from real clinical data and the proposed approach is compared with analyses made with the commonly used parametric IRT graded response model. Practical implications of the proposed methodology are provided.

Ort, förlag, år, upplaga, sidor
Sage Publications, 2026
Nyckelord
surprisal, information manifold, scope, scale information, score index, entropy, mutual entropy
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-242839 (URN)10.3102/10769986251314833 (DOI)001429984000001 ()2-s2.0-86000757638 (Scopus ID)
Forskningsfinansiär
Vetenskapsrådet, 2022-02046Marianne och Marcus Wallenbergs Stiftelse, MMW 2019.0129
Tillgänglig från: 2025-08-08 Skapad: 2025-08-08 Senast uppdaterad: 2026-06-03Bibliografiskt granskad
Ramsay, J. O., Li, J., Wallmark, J. & Wiberg, M. (2025). An information manifold perspective for analyzing test data. Applied psychological measurement, 49(3), 90-108
Öppna denna publikation i ny flik eller fönster >>An information manifold perspective for analyzing test data
2025 (Engelska)Ingår i: Applied psychological measurement, ISSN 0146-6216, E-ISSN 1552-3497, Vol. 49, nr 3, s. 90-108Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Modifications of current psychometric models for analyzing test data are proposed that produce an additive scale measure of information. This information measure is a one-dimensional space curve or curved surface manifold that is invariant across varying manifold indexing systems. The arc length along a curve manifold is used as it is an additive metric having a defined zero and a version of the bit as a unit. This property, referred to here as the scope of the test or an item, facilitates the evaluation of graphs and numerical summaries. The measurement power of the test is defined by the length of the manifold, and the performance or experiential level of a person by a position along the curve. In this study, we also use all information from the items including the information from the distractors. Test data from a large-scale college admissions test are used to illustrate the test information manifold perspective and to compare it with the well-known item response theory nominal model. It is illustrated that the use of information theory opens a vista of new ways of assessing item performance and inter-item dependency, as well as test takers' knowledge.

Ort, förlag, år, upplaga, sidor
Sage Publications, 2025
Nyckelord
entropy, expected sum score, nominal model, scope, score index, spline functions, surprisal, test information, TestGardener
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-233717 (URN)10.1177/01466216241310600 (DOI)001380542200001 ()39713764 (PubMedID)2-s2.0-105001648191 (Scopus ID)
Forskningsfinansiär
Wallenbergstiftelserna, MMW 2019.012
Tillgänglig från: 2025-01-09 Skapad: 2025-01-09 Senast uppdaterad: 2025-04-29Bibliografiskt granskad
Wallmark, J. (2025). Extensions and applications of item response theory. (Doctoral dissertation). Umeå: Umeå University
Öppna denna publikation i ny flik eller fönster >>Extensions and applications of item response theory
2025 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Alternativ titel[sv]
Vidareutveckling och tillämpningar av item response theory
Abstract [en]

This doctoral thesis focuses on Item Response Theory (IRT), a statistical method widely used in fields such as education and psychology to analyze response patterns on tests and surveys. In practice, IRT models are estimated using collected test data, which allows researchers to assess both how effectively each item measures the underlying trait—such as subject knowledge or personality characteristics—that the test aims to evaluate, and to estimate each individual's level of that trait. Unlike traditional methods that simply sum predetermined item scores, IRT accounts for the difficulty of each item and its ability to measure the intended trait.

The thesis consists of four research articles, each addressing different aspects of IRT and its applications. The first article focuses on test equating, ensuring that scores from different versions of a test are comparable. Equating methods with and without IRT are compared using simulations to explore the advantages and disadvantages of incorporating IRT into the kernel equating framework. The second and third articles introduce and compare different types of IRT models. Through simulations and real test data examples, these studies demonstrate that more flexible models can better capture the true relationships between test responses and the underlying traits being measured.

Finally, the IRTorch Python package is presented in the fourth study. IRTorch supports various IRT models and estimation methods and can be used to analyze data from different types of tests and surveys. In summary, the thesis demonstrates how IRT-based equating methods can serve as an alternative to traditional equating methods, how more flexible IRT models can improve the precision of test results, and how user-friendly software can make advanced statistical models accessible to a wider audience.

Ort, förlag, år, upplaga, sidor
Umeå: Umeå University, 2025. s. 25
Serie
Statistical studies, ISSN 1100-8989 ; 60
Nyckelord
Machine learning, Autoencoders, Item response theory, psychometrics, Test equating, Statistical software, Educational assessment, Latent variable modelling
Nationell ämneskategori
Sannolikhetsteori och statistik
Forskningsämne
statistik
Identifikatorer
urn:nbn:se:umu:diva-233351 (URN)978-91-8070-572-1 (ISBN)978-91-8070-571-4 (ISBN)
Disputation
2025-02-07, HUM.D.220 (Hjortronlandet), Humanisthuset, Umeå university, Umeå, 09:00 (Engelska)
Opponent
Handledare
Tillgänglig från: 2025-01-08 Skapad: 2024-12-31 Senast uppdaterad: 2025-01-08Bibliografiskt granskad
Wallmark, J., Ramsay, J. O., Li, J. & Wiberg, M. (2024). Analyzing polytomous test data: a comparison between an information-based IRT model and the generalized partial credit model. Journal of educational and behavioral statistics, 49(5), 753-779
Öppna denna publikation i ny flik eller fönster >>Analyzing polytomous test data: a comparison between an information-based IRT model and the generalized partial credit model
2024 (Engelska)Ingår i: Journal of educational and behavioral statistics, ISSN 1076-9986, E-ISSN 1935-1054, Vol. 49, nr 5, s. 753-779Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Item response theory (IRT) models the relationship between the possible scores on a test item against a test taker’s attainment of the latent trait that the item is intended to measure. In this study, we compare two models for tests with polytomously scored items: the optimal scoring (OS) model, a nonparametric IRT model based on the principles of information theory, and the generalized partial credit (GPC) model, a widely used parametric alternative. We evaluate these models using both simulated and real test data. In the real data examples, the OS model demonstrates superior model fit compared to the GPC model across all analyzed datasets. In our simulation study, the OS model outperforms the GPC model in terms of bias, but at the cost of larger standard errors for the probabilities along the estimated item response functions. Furthermore, we illustrate how surprisal arc length, an IRT scale invariant measure of ability with metric properties, can be used to put scores from vastly different types of IRT models on a common scale. We also demonstrate how arc length can be a viable alternative to sum scores for scoring test takers.

Ort, förlag, år, upplaga, sidor
Sage Publications, 2024
Nyckelord
item characteristic curves, item response theory, nonparametric IRT, simulation
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-216867 (URN)10.3102/10769986231207879 (DOI)001098507600001 ()2-s2.0-85176273777 (Scopus ID)
Forskningsfinansiär
Marianne och Marcus Wallenbergs Stiftelse, 2019.0129
Tillgänglig från: 2023-12-12 Skapad: 2023-12-12 Senast uppdaterad: 2026-06-03Bibliografiskt granskad
Wallmark, J., Josefsson, M. & Wiberg, M. (2023). Efficiency analysis of item response theory kernel equating for mixed-format tests. Applied psychological measurement, 47(7-8), 496-512
Öppna denna publikation i ny flik eller fönster >>Efficiency analysis of item response theory kernel equating for mixed-format tests
2023 (Engelska)Ingår i: Applied psychological measurement, ISSN 0146-6216, E-ISSN 1552-3497, Vol. 47, nr 7-8, s. 496-512Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

This study aims to evaluate the performance of Item Response Theory (IRT) kernel equating in the context of mixed-format tests by comparing it to IRT observed score equating and kernel equating with log-linear presmoothing. Comparisons were made through both simulations and real data applications, under both equivalent groups (EG) and non-equivalent groups with anchor test (NEAT) sampling designs. To prevent bias towards IRT methods, data were simulated with and without the use of IRT models. The results suggest that the difference between IRT kernel equating and IRT observed score equating is minimal, both in terms of the equated scores and their standard errors. The application of IRT models for presmoothing yielded smaller standard error of equating than the log-linear presmoothing approach. When test data were generated using IRT models, IRT-based methods proved less biased than log-linear kernel equating. However, when data were simulated without IRT models, log-linear kernel equating showed less bias. Overall, IRT kernel equating shows great promise when equating mixed-format tests.

Ort, förlag, år, upplaga, sidor
Sage Publications, 2023
Nyckelord
item response theory, kernel equating, log-linear models, presmoothing, simulation
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-215929 (URN)10.1177/01466216231209757 (DOI)001087283200001 ()2-s2.0-85174542085 (Scopus ID)
Forskningsfinansiär
Marianne och Marcus Wallenbergs Stiftelse, 2019.0129
Tillgänglig från: 2023-11-02 Skapad: 2023-11-02 Senast uppdaterad: 2025-04-24Bibliografiskt granskad
Wallmark, J., Josefsson, M. & Wiberg, M. (2023). Kernel equating presmoothing methods: an empirical study with mixed-format test forms. In: Marie Wiberg; Dylan Molenaar; Jorge González; Jee-Seon Kim; Heungsun Hwang (Ed.), Quantitative psychology: The 87th annual meeting of the psychometric society, Bologna, Italy, 2022. Paper presented at 87th Annual Meeting of the Psychometric Society, IMPS 2022, Bologna, Italy, July 11-15, 2022 (pp. 49-59). Springer
Öppna denna publikation i ny flik eller fönster >>Kernel equating presmoothing methods: an empirical study with mixed-format test forms
2023 (Engelska)Ingår i: Quantitative psychology: The 87th annual meeting of the psychometric society, Bologna, Italy, 2022 / [ed] Marie Wiberg; Dylan Molenaar; Jorge González; Jee-Seon Kim; Heungsun Hwang, Springer, 2023, s. 49-59Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

When equating test forms, it is common to presmooth the test score distributions before conducting the equating. In this study, the log-linear and item response theory (IRT) presmoothing methods were compared when equating mixed-format test forms using kernel equating. Test forms from two different high-stakes tests were equated: The Swedish national test in mathematics, using the equivalent group sampling design, and the verbal part of the Swedish SAT test, using the nonequivalent groups with anchor test sampling design. In both cases, the analytical equating standard errors were lower for high and low performing test takers when using IRT presmoothing compared to log-linear presmoothing. Both presmoothing methods resulted in reasonable equated curves. As no true equating transformation is known in a practical setting, using IRT models for presmoothing appears to be a viable alternative to log-linear models when equating mixed-format tests such as the Swedish SAT.

Ort, förlag, år, upplaga, sidor
Springer, 2023
Serie
Springer Proceedings in Mathematics & Statistics, ISSN 2194-1009, E-ISSN 2194-1017 ; 422
Nyckelord
Item response theory, Kernel equating, Presmoothing
Nationell ämneskategori
Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:umu:diva-212317 (URN)10.1007/978-3-031-27781-8_5 (DOI)001310049800005 ()2-s2.0-85164740918 (Scopus ID)9783031277801 (ISBN)
Konferens
87th Annual Meeting of the Psychometric Society, IMPS 2022, Bologna, Italy, July 11-15, 2022
Tillgänglig från: 2023-07-25 Skapad: 2023-07-25 Senast uppdaterad: 2025-04-24Bibliografiskt granskad
Wallmark, J., Josefsson, M. & Wiberg, M.Introducing flexible monotone multiple choice item response theory models and bit scales.
Öppna denna publikation i ny flik eller fönster >>Introducing flexible monotone multiple choice item response theory models and bit scales
(Engelska)Manuskript (preprint) (Övrigt vetenskapligt)
Abstract [en]

Item Response Theory (IRT) is a powerful statistical approach for evaluating test items and determining test taker abilities through response analysis. An IRT model that better fits the data leads to more accurate latent trait estimates. In this study, we present a new model for multiple choice data, the monotone multiple choice (MMC) model, which we fit using autoencoders. Using both simulated scenarios and real data from the Swedish Scholastic Aptitude Test, we demonstrate empirically that the MMC model outperforms the traditional nominal response IRT model in terms of fit. Furthermore, we illustrate how the latent trait scale from any fitted IRT model can be transformed into a ratio scale, aiding in score interpretation and making it easier to compare different types of IRT models. We refer to these new scales as bit scales. Bit scales are especially useful for models for which minimal or no assumptions are made for the latent trait scale distributions, such as for the autoencoder fitted models in this study.

Nyckelord
Item response theory, Neural networks, autoencoders
Nationell ämneskategori
Sannolikhetsteori och statistik
Forskningsämne
statistik; data science
Identifikatorer
urn:nbn:se:umu:diva-233350 (URN)10.48550/arXiv.2410.01480 (DOI)
Forskningsfinansiär
Wallenbergstiftelserna, 2022-02046
Tillgänglig från: 2024-12-31 Skapad: 2024-12-31 Senast uppdaterad: 2025-01-02Bibliografiskt granskad
Wallmark, J.IRTorch: an Item Response Theory Python package.
Öppna denna publikation i ny flik eller fönster >>IRTorch: an Item Response Theory Python package
(Engelska)Manuskript (preprint) (Övrigt vetenskapligt)
Abstract [en]

Item Response Theory (IRT) is a statistical framework used to model the relationshipbetween latent traits (such as abilities or personality traits) and responses to items meantto assess those traits. In this article, we introduce the IRTorch Python package for fittingand evaluating IRT models. The package utilizes PyTorch for parameter optimizationand GPU support. It supports a diverse range of unidimensional and multidimensionalIRT models, both parametric and semiparametric. IRTorch also emphasizes the arbitrarynature of the latent variable scale, which is implicitly assumed and often ignored in otherIRT software. The package provides a flexible framework to implement custom models,scale transformations, and fitting algorithms. We illustrate some of the package’s featuresthrough several examples, including fitting traditional IRT models, using autoencoders forfitting IRT models, and using the bit scale transformation to give a unit of measurementto the latent trait scale.

Nyckelord
IRT, Python, PyTorch, model estimation, autoencoders
Nationell ämneskategori
Sannolikhetsteori och statistik
Forskningsämne
statistik; pedagogik; psykologi
Identifikatorer
urn:nbn:se:umu:diva-233349 (URN)
Forskningsfinansiär
Vetenskapsrådet, 022-02046
Tillgänglig från: 2024-12-31 Skapad: 2024-12-31 Senast uppdaterad: 2025-01-02Bibliografiskt granskad
Organisationer
Identifikatorer
ORCID-id: ORCID iD iconorcid.org/0000-0001-7573-0671

Sök vidare i DiVA

Visa alla publikationer