Umeå University's logo

umu.sePublications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 10) Show all publications
Wallmark, J., Wiberg, M. & Eriksson, M. (2026). Developing and validating a frailty score based on patient-reported outcome 3 months after stroke: a Riksstroke-based study. PLOS ONE, 21(2), Article ID e0343249.
Open this publication in new window or tab >>Developing and validating a frailty score based on patient-reported outcome 3 months after stroke: a Riksstroke-based study
2026 (English)In: PLOS ONE, E-ISSN 1932-6203, Vol. 21, no 2, article id e0343249Article in journal (Refereed) Published
Abstract [en]

BACKGROUND: Frailty is common after stroke and linked to poor outcomes, but many measures are clinician-rated, time-consuming, and not suited to patient-reported data. To address these issues, we developed and validated a frailty score from the Swedish Stroke Register (Riksstroke) three-month follow-up questionnaire.

METHODS: We analyzed responses from 19,470 stroke survivors to nine patient-reported items covering function, mood, fatigue, pain and general health, in the 2021-2022 Riksstroke questionnaire. Dimensionality was assessed with Mokken Scale Analysis and exploratory factor analysis. Item response theory (IRT) was used for score computation. Competing graded response IRT models (unidimensional, correlated-factor, bifactor) were compared, and measurement fairness was examined using differential item functioning (DIF) across age, sex, and education. Prognostic validity was tested with Kaplan-Meier curves and Cox regression for all-cause mortality.

RESULTS: From the Mokken Scale Analysis, all items met scalability criteria. Factor analysis suggested two correlated interpretable facets (Physical Functioning; Well-being/Mental Health). A bifactor IRT model provided the best fit to the data, comprising a general frailty dimension while addressing the strong correlation between the facets. DIF was minimal for sex and education, with modest age-related effects. Higher frailty scores were associated with increased mortality in adjusted Cox models and Kaplan-Meier curves. Tools for computing frailty scores are available at https://github.com/joakimwallmark/frailty-irt-scores.

CONCLUSIONS: A robust, fair, and prognostically meaningful frailty score can be derived from patient-reported items in Riksstroke. More broadly, the study demonstrates how routinely collected patient-reported outcome measures can be leveraged to build scalable frailty scores, offering efficient cost-effective tools for monitoring outcome and guiding quality improvement in stroke care.

Place, publisher, year, edition, pages
Public Library of Science (PLoS), 2026
National Category
Medical Biostatistics
Research subject
Statistics; Epidemiology
Identifiers
urn:nbn:se:umu:diva-250646 (URN)10.1371/journal.pone.0343249 (DOI)001695758500001 ()41719319 (PubMedID)2-s2.0-105030698386 (Scopus ID)
Funder
Swedish Research Council, 2022-02046Swedish Research Council, 2024-02846
Available from: 2026-03-04 Created: 2026-03-04 Last updated: 2026-03-13Bibliographically approved
Wallmark, J. & Wiberg, M. (2026). The bit scale: a metric score scale for unidimensional item response theory models. Psychometrika, 91(1), 30-46
Open this publication in new window or tab >>The bit scale: a metric score scale for unidimensional item response theory models
2026 (English)In: Psychometrika, ISSN 0033-3123, E-ISSN 1860-0980, Vol. 91, no 1, p. 30-46Article in journal (Refereed) Published
Abstract [en]

In Item Response Theory (IRT), the conventional latent trait scale (θ) is inherently arbitrary, lacking a fixed unit or origin and often tied to specific population distributional assumptions (e.g., standard normal). This limits the direct comparability and interpretability of scores across different tests, populations, or model estimation methods. This paper introduces the “bit scale,” a novel metric transformation for unidimensional IRT scores derived from fundamental principles of information theory, specifically surprisal and entropy. Bit scores are anchored to the properties of the test items rather than the test-taker population. This item-based anchoring ensures the scale’s invariance to population assumptions and provides a consistent metric for comparing latent trait levels. We illustrate the utility of the bit scale through empirical examples: demonstrating consistent scoring when fitting models with different θ scale assumptions, and using anchor items to directly link scores from different test administrations. A simulation study confirms the desirable statistical properties (low bias, accurate standard errors) of Maximum Likelihood estimated bit scores and their robustness to extreme scores. The bit scale offers a theoretically grounded, interpretable, and comparable metric for reporting and analyzing IRT-based assessment results. Software implementations in R (bitscale) and Python (IRTorch) are available and practical implications are discussed.

Place, publisher, year, edition, pages
Cambridge University Press, 2026
Keywords
bit scale, bit scores, information, rescaling, scale transformation
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-248215 (URN)10.1017/psy.2025.10071 (DOI)001658889300001 ()41367367 (PubMedID)2-s2.0-105025666005 (Scopus ID)
Funder
Marianne and Marcus Wallenberg Foundation, 2019-0129Swedish Research Council, 2022-02046
Available from: 2026-01-09 Created: 2026-01-09 Last updated: 2026-05-21Bibliographically approved
Ramsay, J. O., Li, J., Bernstein, C. N., Marrie, R. A., Wallmark, J. & Wiberg, M. (2026). Using the information metric to analyze clinical rating scales. Journal of educational and behavioral statistics, 51(2), 395-418
Open this publication in new window or tab >>Using the information metric to analyze clinical rating scales
Show others...
2026 (English)In: Journal of educational and behavioral statistics, ISSN 1076-9986, E-ISSN 1935-1054, Vol. 51, no 2, p. 395-418Article in journal (Refereed) Published
Abstract [en]

A rating scale is a set of categories designed to obtain information about a quantitative or a qualitative attribute. Item response theory (IRT) proposes that a probability function over a single latent variable represents the overall attribute evolution that the scale is designed to assess. Here we utilize an information theory approach to IRT to analyze rating scale data. The proposed IRT analyses, based on surprisal, offer new tools for assessing raters, rated items, and the whole rating scale. The information transformation from probability to surprisal is a new lens from which to view choice data and is an important augmentation of probability-based IRT. It also offers new graphical tools to measure the amount of information captured by an item in an additive metric, and to measure covariation among items using mutual information. The proposed methodology is illustrated using two scales from real clinical data and the proposed approach is compared with analyses made with the commonly used parametric IRT graded response model. Practical implications of the proposed methodology are provided.

Place, publisher, year, edition, pages
Sage Publications, 2026
Keywords
surprisal, information manifold, scope, scale information, score index, entropy, mutual entropy
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-242839 (URN)10.3102/10769986251314833 (DOI)001429984000001 ()2-s2.0-86000757638 (Scopus ID)
Funder
Swedish Research Council, 2022-02046Marianne and Marcus Wallenberg Foundation, MMW 2019.0129
Available from: 2025-08-08 Created: 2025-08-08 Last updated: 2026-06-03Bibliographically approved
Ramsay, J. O., Li, J., Wallmark, J. & Wiberg, M. (2025). An information manifold perspective for analyzing test data. Applied psychological measurement, 49(3), 90-108
Open this publication in new window or tab >>An information manifold perspective for analyzing test data
2025 (English)In: Applied psychological measurement, ISSN 0146-6216, E-ISSN 1552-3497, Vol. 49, no 3, p. 90-108Article in journal (Refereed) Published
Abstract [en]

Modifications of current psychometric models for analyzing test data are proposed that produce an additive scale measure of information. This information measure is a one-dimensional space curve or curved surface manifold that is invariant across varying manifold indexing systems. The arc length along a curve manifold is used as it is an additive metric having a defined zero and a version of the bit as a unit. This property, referred to here as the scope of the test or an item, facilitates the evaluation of graphs and numerical summaries. The measurement power of the test is defined by the length of the manifold, and the performance or experiential level of a person by a position along the curve. In this study, we also use all information from the items including the information from the distractors. Test data from a large-scale college admissions test are used to illustrate the test information manifold perspective and to compare it with the well-known item response theory nominal model. It is illustrated that the use of information theory opens a vista of new ways of assessing item performance and inter-item dependency, as well as test takers' knowledge.

Place, publisher, year, edition, pages
Sage Publications, 2025
Keywords
entropy, expected sum score, nominal model, scope, score index, spline functions, surprisal, test information, TestGardener
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-233717 (URN)10.1177/01466216241310600 (DOI)001380542200001 ()39713764 (PubMedID)2-s2.0-105001648191 (Scopus ID)
Funder
Wallenberg Foundations, MMW 2019.012
Available from: 2025-01-09 Created: 2025-01-09 Last updated: 2025-04-29Bibliographically approved
Wallmark, J. (2025). Extensions and applications of item response theory. (Doctoral dissertation). Umeå: Umeå University
Open this publication in new window or tab >>Extensions and applications of item response theory
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Alternative title[sv]
Vidareutveckling och tillämpningar av item response theory
Abstract [en]

This doctoral thesis focuses on Item Response Theory (IRT), a statistical method widely used in fields such as education and psychology to analyze response patterns on tests and surveys. In practice, IRT models are estimated using collected test data, which allows researchers to assess both how effectively each item measures the underlying trait—such as subject knowledge or personality characteristics—that the test aims to evaluate, and to estimate each individual's level of that trait. Unlike traditional methods that simply sum predetermined item scores, IRT accounts for the difficulty of each item and its ability to measure the intended trait.

The thesis consists of four research articles, each addressing different aspects of IRT and its applications. The first article focuses on test equating, ensuring that scores from different versions of a test are comparable. Equating methods with and without IRT are compared using simulations to explore the advantages and disadvantages of incorporating IRT into the kernel equating framework. The second and third articles introduce and compare different types of IRT models. Through simulations and real test data examples, these studies demonstrate that more flexible models can better capture the true relationships between test responses and the underlying traits being measured.

Finally, the IRTorch Python package is presented in the fourth study. IRTorch supports various IRT models and estimation methods and can be used to analyze data from different types of tests and surveys. In summary, the thesis demonstrates how IRT-based equating methods can serve as an alternative to traditional equating methods, how more flexible IRT models can improve the precision of test results, and how user-friendly software can make advanced statistical models accessible to a wider audience.

Place, publisher, year, edition, pages
Umeå: Umeå University, 2025. p. 25
Series
Statistical studies, ISSN 1100-8989 ; 60
Keywords
Machine learning, Autoencoders, Item response theory, psychometrics, Test equating, Statistical software, Educational assessment, Latent variable modelling
National Category
Probability Theory and Statistics
Research subject
Statistics
Identifiers
urn:nbn:se:umu:diva-233351 (URN)978-91-8070-572-1 (ISBN)978-91-8070-571-4 (ISBN)
Public defence
2025-02-07, HUM.D.220 (Hjortronlandet), Humanisthuset, Umeå university, Umeå, 09:00 (English)
Opponent
Supervisors
Available from: 2025-01-08 Created: 2024-12-31 Last updated: 2025-01-08Bibliographically approved
Wallmark, J., Ramsay, J. O., Li, J. & Wiberg, M. (2024). Analyzing polytomous test data: a comparison between an information-based IRT model and the generalized partial credit model. Journal of educational and behavioral statistics, 49(5), 753-779
Open this publication in new window or tab >>Analyzing polytomous test data: a comparison between an information-based IRT model and the generalized partial credit model
2024 (English)In: Journal of educational and behavioral statistics, ISSN 1076-9986, E-ISSN 1935-1054, Vol. 49, no 5, p. 753-779Article in journal (Refereed) Published
Abstract [en]

Item response theory (IRT) models the relationship between the possible scores on a test item against a test taker’s attainment of the latent trait that the item is intended to measure. In this study, we compare two models for tests with polytomously scored items: the optimal scoring (OS) model, a nonparametric IRT model based on the principles of information theory, and the generalized partial credit (GPC) model, a widely used parametric alternative. We evaluate these models using both simulated and real test data. In the real data examples, the OS model demonstrates superior model fit compared to the GPC model across all analyzed datasets. In our simulation study, the OS model outperforms the GPC model in terms of bias, but at the cost of larger standard errors for the probabilities along the estimated item response functions. Furthermore, we illustrate how surprisal arc length, an IRT scale invariant measure of ability with metric properties, can be used to put scores from vastly different types of IRT models on a common scale. We also demonstrate how arc length can be a viable alternative to sum scores for scoring test takers.

Place, publisher, year, edition, pages
Sage Publications, 2024
Keywords
item characteristic curves, item response theory, nonparametric IRT, simulation
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-216867 (URN)10.3102/10769986231207879 (DOI)001098507600001 ()2-s2.0-85176273777 (Scopus ID)
Funder
Marianne and Marcus Wallenberg Foundation, 2019.0129
Available from: 2023-12-12 Created: 2023-12-12 Last updated: 2026-06-03Bibliographically approved
Wallmark, J., Josefsson, M. & Wiberg, M. (2023). Efficiency analysis of item response theory kernel equating for mixed-format tests. Applied psychological measurement, 47(7-8), 496-512
Open this publication in new window or tab >>Efficiency analysis of item response theory kernel equating for mixed-format tests
2023 (English)In: Applied psychological measurement, ISSN 0146-6216, E-ISSN 1552-3497, Vol. 47, no 7-8, p. 496-512Article in journal (Refereed) Published
Abstract [en]

This study aims to evaluate the performance of Item Response Theory (IRT) kernel equating in the context of mixed-format tests by comparing it to IRT observed score equating and kernel equating with log-linear presmoothing. Comparisons were made through both simulations and real data applications, under both equivalent groups (EG) and non-equivalent groups with anchor test (NEAT) sampling designs. To prevent bias towards IRT methods, data were simulated with and without the use of IRT models. The results suggest that the difference between IRT kernel equating and IRT observed score equating is minimal, both in terms of the equated scores and their standard errors. The application of IRT models for presmoothing yielded smaller standard error of equating than the log-linear presmoothing approach. When test data were generated using IRT models, IRT-based methods proved less biased than log-linear kernel equating. However, when data were simulated without IRT models, log-linear kernel equating showed less bias. Overall, IRT kernel equating shows great promise when equating mixed-format tests.

Place, publisher, year, edition, pages
Sage Publications, 2023
Keywords
item response theory, kernel equating, log-linear models, presmoothing, simulation
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-215929 (URN)10.1177/01466216231209757 (DOI)001087283200001 ()2-s2.0-85174542085 (Scopus ID)
Funder
Marianne and Marcus Wallenberg Foundation, 2019.0129
Available from: 2023-11-02 Created: 2023-11-02 Last updated: 2025-04-24Bibliographically approved
Wallmark, J., Josefsson, M. & Wiberg, M. (2023). Kernel equating presmoothing methods: an empirical study with mixed-format test forms. In: Marie Wiberg; Dylan Molenaar; Jorge González; Jee-Seon Kim; Heungsun Hwang (Ed.), Quantitative psychology: The 87th annual meeting of the psychometric society, Bologna, Italy, 2022. Paper presented at 87th Annual Meeting of the Psychometric Society, IMPS 2022, Bologna, Italy, July 11-15, 2022 (pp. 49-59). Springer
Open this publication in new window or tab >>Kernel equating presmoothing methods: an empirical study with mixed-format test forms
2023 (English)In: Quantitative psychology: The 87th annual meeting of the psychometric society, Bologna, Italy, 2022 / [ed] Marie Wiberg; Dylan Molenaar; Jorge González; Jee-Seon Kim; Heungsun Hwang, Springer, 2023, p. 49-59Conference paper, Published paper (Refereed)
Abstract [en]

When equating test forms, it is common to presmooth the test score distributions before conducting the equating. In this study, the log-linear and item response theory (IRT) presmoothing methods were compared when equating mixed-format test forms using kernel equating. Test forms from two different high-stakes tests were equated: The Swedish national test in mathematics, using the equivalent group sampling design, and the verbal part of the Swedish SAT test, using the nonequivalent groups with anchor test sampling design. In both cases, the analytical equating standard errors were lower for high and low performing test takers when using IRT presmoothing compared to log-linear presmoothing. Both presmoothing methods resulted in reasonable equated curves. As no true equating transformation is known in a practical setting, using IRT models for presmoothing appears to be a viable alternative to log-linear models when equating mixed-format tests such as the Swedish SAT.

Place, publisher, year, edition, pages
Springer, 2023
Series
Springer Proceedings in Mathematics & Statistics, ISSN 2194-1009, E-ISSN 2194-1017 ; 422
Keywords
Item response theory, Kernel equating, Presmoothing
National Category
Probability Theory and Statistics
Identifiers
urn:nbn:se:umu:diva-212317 (URN)10.1007/978-3-031-27781-8_5 (DOI)001310049800005 ()2-s2.0-85164740918 (Scopus ID)9783031277801 (ISBN)
Conference
87th Annual Meeting of the Psychometric Society, IMPS 2022, Bologna, Italy, July 11-15, 2022
Available from: 2023-07-25 Created: 2023-07-25 Last updated: 2025-04-24Bibliographically approved
Wallmark, J., Josefsson, M. & Wiberg, M.Introducing flexible monotone multiple choice item response theory models and bit scales.
Open this publication in new window or tab >>Introducing flexible monotone multiple choice item response theory models and bit scales
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Item Response Theory (IRT) is a powerful statistical approach for evaluating test items and determining test taker abilities through response analysis. An IRT model that better fits the data leads to more accurate latent trait estimates. In this study, we present a new model for multiple choice data, the monotone multiple choice (MMC) model, which we fit using autoencoders. Using both simulated scenarios and real data from the Swedish Scholastic Aptitude Test, we demonstrate empirically that the MMC model outperforms the traditional nominal response IRT model in terms of fit. Furthermore, we illustrate how the latent trait scale from any fitted IRT model can be transformed into a ratio scale, aiding in score interpretation and making it easier to compare different types of IRT models. We refer to these new scales as bit scales. Bit scales are especially useful for models for which minimal or no assumptions are made for the latent trait scale distributions, such as for the autoencoder fitted models in this study.

Keywords
Item response theory, Neural networks, autoencoders
National Category
Probability Theory and Statistics
Research subject
Statistics; data science
Identifiers
urn:nbn:se:umu:diva-233350 (URN)10.48550/arXiv.2410.01480 (DOI)
Funder
Wallenberg Foundations, 2022-02046
Available from: 2024-12-31 Created: 2024-12-31 Last updated: 2025-01-02Bibliographically approved
Wallmark, J.IRTorch: an Item Response Theory Python package.
Open this publication in new window or tab >>IRTorch: an Item Response Theory Python package
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Item Response Theory (IRT) is a statistical framework used to model the relationshipbetween latent traits (such as abilities or personality traits) and responses to items meantto assess those traits. In this article, we introduce the IRTorch Python package for fittingand evaluating IRT models. The package utilizes PyTorch for parameter optimizationand GPU support. It supports a diverse range of unidimensional and multidimensionalIRT models, both parametric and semiparametric. IRTorch also emphasizes the arbitrarynature of the latent variable scale, which is implicitly assumed and often ignored in otherIRT software. The package provides a flexible framework to implement custom models,scale transformations, and fitting algorithms. We illustrate some of the package’s featuresthrough several examples, including fitting traditional IRT models, using autoencoders forfitting IRT models, and using the bit scale transformation to give a unit of measurementto the latent trait scale.

Keywords
IRT, Python, PyTorch, model estimation, autoencoders
National Category
Probability Theory and Statistics
Research subject
Statistics; education; Psychology
Identifiers
urn:nbn:se:umu:diva-233349 (URN)
Funder
Swedish Research Council, 022-02046
Available from: 2024-12-31 Created: 2024-12-31 Last updated: 2025-01-02Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-7573-0671

Search in DiVA

Show all publications