Should You Compare Wearable Data Against the PSQI, and How Do You Interpret Discrepancies?

It is one of the most common questions we hear from teams designing real-world wearable sleep studies: you have a device producing nightly estimates of total sleep time, sleep efficiency, and wake after sleep onset, and you have a validated questionnaire like the Pittsburgh Sleep Quality Index (PSQI) sitting in your enrollment battery. The instinct is to check one against the other. When the two disagree, the reflex is to ask which one is wrong.

The short answer: do not treat the PSQI as a criterion standard for validating a wearable. The two instruments measure different constructs, so weak agreement is the expected result, not a sign of device failure. But that does not mean you should leave the PSQI out. Used correctly, the gap between subjective report and objective measurement is one of the most informative signals in your dataset. This post explains the mechanism behind the mismatch, the evidence for it, and a staged framework for interpreting discrepancies.

The mechanism: two instruments, two constructs

The PSQI is a 19-item self-report questionnaire that produces seven component scores (subjective sleep quality, sleep latency, sleep duration, habitual sleep efficiency, sleep disturbances, use of sleep medication, and daytime dysfunction) summed into a global score from 0 to 21. A global score above 5 indicates poor sleep quality, with sensitivity of 89.6 percent and specificity of 86.5 percent in the original validation.1 Critically, the PSQI asks respondents to summarize their sleep over the past month, and the instrument is validated only for that one-month recall window.

So the PSQI captures a retrospective, trait-level, subjective appraisal of habitual sleep. It is closer to sleep satisfaction than to any single night of physiology. A wearable, by contrast, produces a prospective, night-level, objective estimate of sleep and wake derived from accelerometry and photoplethysmography. These are categorically different quantities. Comparing them and expecting tight agreement is a bit like comparing a person’s month-end recollection of how well they ate against a food scale reading from last Tuesday’s dinner. Both are legitimate, but they answer different questions on different timescales.

This distinction was recognized in the original PSQI paper itself, which discussed the lack of agreement between the global score and polysomnographic variables and attributed it to the instrument requesting a global estimate spanning a month rather than a night-specific measurement.

The evidence: PSQI and objective metrics agree weakly

The empirical literature is remarkably consistent. When PSQI-derived sleep parameters are compared against actigraphy, the correlations are weak and frequently non-significant.

  • In 71 healthy midlife women, PSQI global score correlated with actigraphy total sleep time at rho = -0.10 (non-significant). At the component level, PSQI sleep efficiency versus actigraphy sleep efficiency reached only r = 0.05, and PSQI sleep onset latency versus actigraphy latency reached r = 0.24, neither significant. The authors concluded the latency and efficiency components were not congruent with their objective homologues.
  • In 63 overweight and obese adults, PSQI-derived total sleep time correlated with actigraphy at r = 0.20 to 0.31, and the average discrepancy exceeded 50 minutes. Participants with a PSQI above 5 showed the largest discrepancies.
  • In 119 healthy women, there was no significant relationship between PSQI ratings and actigraphy-measured sleep duration, efficiency, or latency, while a prospective sleep diary tracked actigraphy well.
  • In 112 non-clinical volunteers, the PSQI global score correlated appreciably with sleep-diary variables and depression scores but with no actigraphic sleep variable. The authors suggested the PSQI may reflect a negative cognitive viewpoint more than actual sleep parameters.
  • In 113 adults, overall correlation between actigraphy total sleep time and PSQI-reported time was r = 0.12 (non-significant), though supplying bedtimes and wake times improved the relationship substantially (r = 0.57).

Where the PSQI does show stronger convergent validity is against other subjective measures. It correlates around r = 0.53 with diary-derived sleep efficiency and r = 0.63 with the Insomnia Severity Index. In other words, the PSQI behaves exactly as a subjective construct should: it clusters with other subjective and affective measures, not with objective sleep. That is by design, not defect.

Why validating a wearable against the PSQI is a category error

Validation asks whether a device accurately measures what it claims to measure. Because the PSQI does not measure per-night objective sleep, it cannot serve as the criterion for whether a wearable measures per-night objective sleep accurately. Using it that way conflates convergent validity across constructs with analytical accuracy. The field’s consensus reference standard is polysomnography for physiological sleep and staging, with actigraphy (itself validated against polysomnography) as the accepted mobile comparator.

The Digital Medicine Society (DiMe) V3 framework makes the roles explicit.Verification confirms the sensor captures the raw signal accurately. Analytical validation confirms the algorithm computes the sleep metric accurately against a physiological reference standard, and self-report has no role here. Clinical validation confirms the metric relates to a meaningful health concept, and this is precisely where self-report instruments like the PSQI legitimately enter, as one dimension of meaningfulness rather than as an accuracy yardstick. The V3+ extension adds usability validation on top of this structure. If you find yourself listing the PSQI as your analytical-validation reference, you have placed a clinical-validation instrument in an analytical-validation slot.

Discrepancies are often the signal, not the noise

The most important reframing for practitioners: the gap between what a person reports and what their physiology shows is a measurable, patterned, clinically meaningful quantity. It has a name, subjective-objective sleep discrepancy (SOSD), also called sleep misperception.

The foundational synthesis documents that people with insomnia characteristically overestimate sleep onset latency and underestimate total sleep time relative to objective measures, a tendency described as ubiquitous although not universal. The direction and magnitude of the discrepancy vary systematically by population. Insomnia with objectively normal sleep duration is associated with cortical and cognitive-emotional arousal and misperception, whereas insomnia with objectively short sleep duration is the more biologically severe phenotype. Recent in-home electroencephalography work found that 44.9 percent of people who subjectively perceived their sleep as sufficient were objectively classified as insufficient, with overestimation of total sleep time growing as objective sleep worsened. The discrepancy carries prognostic weight too: it has been associated with mortality in older adults and predicts response to cognitive behavioral therapy for insomnia.

There is an important interaction here that makes discrepancies larger than you might expect. Wearables and actigraphy tend to overestimate total sleep time and efficiency and underestimate wake, because motionless wakefulness is easily misclassified as sleep. This is the mirror image of insomnia misperception. So a device may report more and better sleep at exactly the moment an insomnia patient reports less and worse sleep, widening the apparent gap from both directions at once.

A five-way framework for interpreting a wearable-PSQI gap

When device and questionnaire disagree, diagnose the cause before drawing a conclusion. There are five candidate explanations, and they are not mutually exclusive.

  1. Construct difference (expected). Subjective month-level quality and objective night-level quantity will diverge. Small to moderate disagreement is the null expectation, not a red flag.
  2. Genuine sleep misperception (signal). If the direction matches known phenotypes, for example an insomnia participant underestimating sleep time and overestimating latency, treat the gap as a clinical signal and quantify it explicitly.
  3. Device error, non-wear, or compliance. Check wear-time logs, missing nights, firmware changes, and known device biases before attributing anything to the participant.
  4. Recall bias in the PSQI. Month-level retrospection is sensitive to mood and salience. Low mood inflates sleep complaints independent of physiology.
  5. Time-window mismatch. The PSQI’s one-month recall rarely aligns with a specific device measurement window. Align the windows before comparing anything.

The practical rule follows directly. If you need to know whether the device is accurate, compare it to polysomnography or validated actigraphy. If you need to know whether the person’s perception diverges from their physiology, compare the device to self-report and call that discrepancy SOSD, not error.

Better subjective comparators than the PSQI

If your goal is night-level comparison against objective data, the PSQI is the wrong subjective instrument because of its monthly recall. The Consensus Sleep Diary is a standardized prospective self-report that is collected night by night, so it aligns temporally with wearable data far better than the PSQI.For insomnia symptom severity, the Insomnia Severity Index provides a 0 to 28 score with a two-week recall and a well-validated cutoff.For calibrated patient-reported outcomes, the PROMIS Sleep Disturbance and Sleep-Related Impairment item banks offer strong measurement properties and discrimination for moderate to severe insomnia.

A useful mental model: reserve the PSQI for trait-level characterization, screening, and covariate adjustment. Use a prospective diary for anything you intend to line up against nightly device output.

Calibrating expectations for the objective side

Before you attribute a discrepancy to a participant’s perception, remember that the objective side has its own error profile. Even against polysomnography, consumer devices show characteristic bias. A multi-device validation of seven consumer trackers found most performed comparably to or better than actigraphy for sleep and wake detection, though sleep-staging performance was more limited. A meta-analysis of 24 studies and 798 participants reported that wrist-worn devices differed significantly from polysomnography, with a total sleep time mean difference of about -17 minutes and a sleep efficiency difference of about -4.7 percent. A recent six-device comparison found all devices detected more than 90 percent of sleep epochs but had low wake specificity (roughly 29 to 52 percent), with four-stage agreement in the fair-to-moderate range. Four-stage classification accuracy across current devices generally falls in the 60 to 75 percent range.

The lesson is not that devices are unreliable. It is that both sides of the comparison carry structured error, so an unexamined gap tells you very little on its own.

A staged decision framework

Design. Decide the question first. If the goal is device accuracy, pre-register polysomnography or polysomnography-validated actigraphy as the analytical-validation reference and follow a standardized performance-testing framework. Do not list the PSQI as a validation criterion. If the goal is sleep-health characterization or an intervention, include the PSQI at baseline for screening and covariate use.

Instrument selection. Add a Consensus Sleep Diary concurrent with device wear for night-level subjective comparison. Reserve the PSQI for trait-level characterization. Add the Insomnia Severity Index if insomnia symptoms are relevant, and PROMIS measures if you need calibrated patient-reported outcome scores. Align recall windows to device windows wherever possible.

Analysis. Report device-versus-reference accuracy (Bland-Altman bias and limits of agreement, ICC, epoch-by-epoch sensitivity, specificity, and kappa) separately from device-versus-self-report discrepancy. Model SOSD explicitly, for example as a difference or ratio index, as an outcome or moderator rather than as error to be minimized.

Interpretation. Apply the five-way framework. Before attributing a gap to the device, rule out non-wear, compliance, and time-window mismatch. Before attributing it to the participant, confirm that the device’s known biases do not already explain the direction and size of the gap.

Where raw signal access changes the picture

Every step of the framework above depends on your ability to separate three things: what the algorithm computed, what the raw physiology actually showed, and what the participant perceived. That separation is only possible when you can reach the underlying signal.

This is the practical reason Centralive is built on raw-signal hardware SDK access. Through the Garmin Health Companion SDK, studies can collect raw beat-to-beat intervals, accelerometry, and per-beat confidence flags directly, and through the Apple SDK they can access device-level signal rather than only summary outputs. When you have the raw signal, you can run your own analytical validation against polysomnography or actigraphy, apply consistent open-source pipelines across participants, and compute discrepancy indices cleanly because you know exactly how the objective estimate was derived. You can also freeze research firmware so an algorithm update does not silently shift your metrics mid-study.

The contrast with an API-only architecture is structural rather than a matter of product preference. A platform such as Oura returns processed outputs only, with no path to the raw signal beneath them. In that setup, when a wearable estimate and a self-report diverge, you cannot decompose the gap. You cannot tell whether the discrepancy reflects genuine sleep misperception, a quirk of the vendor’s scoring algorithm, or a period of poor signal quality, because the intermediate layers are sealed. For a study whose scientific value lives in interpreting discrepancies, that opacity removes the very thing you came to measure. Raw-signal access is what turns a puzzling gap into an analyzable one.

Bottom line

Do not validate a wearable against the PSQI. They measure different constructs, and weak agreement is expected rather than diagnostic of device error. Do include a subjective measure, ideally a prospective diary aligned to your device window, and treat the subjective-objective discrepancy as a meaningful outcome in its own right. Interpret every gap through the five-way framework, keep analytical validation anchored to polysomnography, and make sure you can reach the raw signal so that when device and self-report disagree, you can actually explain why.

References

  1. Buysse DJ, Reynolds CF, Monk TH, Berman SR, Kupfer DJ. The Pittsburgh Sleep Quality Index: a new instrument for psychiatric practice and research. Psychiatry Research. 1989;28(2):193-213. PMID 2748771. https://pubmed.ncbi.nlm.nih.gov/2748771/
  2. Zak RS, Zitser J, Jones HJ, Gilliss CL, Lee KA. Sleep self-report and actigraphy measures in healthy midlife women: validity of the Pittsburgh Sleep Quality Index. Journal of Women’s Health. 2022;31(7):965-973. DOI 10.1089/jwh.2021.0328. https://doi.org/10.1089/jwh.2021.0328
  3. O’Brien E, Hart C, Wing RR. Discrepancies between self-reported usual sleep duration and objective measures of total sleep time in treatment-seeking overweight and obese individuals. Behavioral Sleep Medicine. 2016;14(5):539-549. DOI 10.1080/15402002.2015.1048447. https://doi.org/10.1080/15402002.2015.1048447
  4. Jackowska M, Ronaldson A, Brown J, Steptoe A. Biological and psychological correlates of self-reported and objective sleep measures. Journal of Psychosomatic Research. 2016;84:52-55. DOI 10.1016/j.jpsychores.2016.03.017. https://doi.org/10.1016/j.jpsychores.2016.03.017
  5. Grandner MA, Kripke DF, Yoon I, Youngstedt SD. Criterion validity of the Pittsburgh Sleep Quality Index: investigation in a non-clinical sample. Sleep and Biological Rhythms. 2006;4(2):129-139. DOI 10.1111/j.1479-8425.2006.00207.x. https://doi.org/10.1111/j.1479-8425.2006.00207.x
  6. Blaxton JM, Bergeman CS, Wang L. Information on bedtimes and wake times improves the relation between self-reported and objective assessments of sleep in adults. Journal of Clinical Sleep Medicine. 2019. DOI 10.5664/jcsm.7888. https://doi.org/10.5664/jcsm.7888
  7. Dietch JR, Taylor DJ, Sethi K, Kelly K, Bramoweth AD, Roane BM. Psychometric evaluation of the PSQI in U.S. college students. Journal of Clinical Sleep Medicine. 2016;12(8):1121-1129. DOI 10.5664/jcsm.6050. https://doi.org/10.5664/jcsm.6050
  8. de Zambotti M, Cellini N, Goldstone A, Colrain IM, Baker FC. Wearable sleep technology in clinical and research settings. Medicine and Science in Sports and Exercise. 2019;51(7):1538-1557. DOI 10.1249/MSS.0000000000001947. https://doi.org/10.1249/MSS.0000000000001947
  9. Menghini L, Cellini N, Goldstone A, Baker FC, de Zambotti M. A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. Sleep. 2021;44(2):zsaa170. DOI 10.1093/sleep/zsaa170. https://doi.org/10.1093/sleep/zsaa170
  10. Goldsack JC, Coravos A, Bakker JP, et al. Verification, analytical validation, and clinical validation (V3): the foundation of determining fit-for-purpose for biometric monitoring technologies (BioMeTs). npj Digital Medicine. 2020;3:55. DOI 10.1038/s41746-020-0260-4. https://doi.org/10.1038/s41746-020-0260-4
  11. Vandendriessche B, et al. V3+ framework to ensure user-centricity and scalability of sensor-based digital health technologies. npj Digital Medicine. 2024;7. DOI 10.1038/s41746-024-01322-2. https://doi.org/10.1038/s41746-024-01322-2 (verify authors, volume, and pages before publication)
  12. Harvey AG, Tang NKY. (Mis)perception of sleep in insomnia: a puzzle and a resolution. Psychological Bulletin. 2012;138(1):77-101. DOI 10.1037/a0025730. https://doi.org/10.1037/a0025730
  13. Fernandez-Mendoza J, Calhoun SL, Bixler EO, et al. Sleep misperception and chronic insomnia in the general population: role of objective sleep duration and psychological profiles. Psychosomatic Medicine. 2011;73(1):88-97. DOI 10.1097/PSY.0b013e3181fe365a. https://doi.org/10.1097/PSY.0b013e3181fe365a
  14. Vgontzas AN, Fernandez-Mendoza J, Liao D, Bixler EO. Insomnia with objective short sleep duration: the most biologically severe phenotype of the disorder. Sleep Medicine Reviews. 2013;17(4):241-254. DOI 10.1016/j.smrv.2012.09.005. https://doi.org/10.1016/j.smrv.2012.09.005
  15. Discrepancies between subjective and objective sleep assessments revealed by in-home electroencephalography during real-world sleep. Proceedings of the National Academy of Sciences. 2025. DOI 10.1073/pnas.2412895121. https://doi.org/10.1073/pnas.2412895121 (verify authors, volume, and pages before publication)
  16. Subjective-objective discrepancy in sleep duration and mortality in older men. Scientific Reports. 2022;12. DOI 10.1038/s41598-022-22065-8. https://doi.org/10.1038/s41598-022-22065-8 (verify authors and article number before publication)
  17. Carney CE, Buysse DJ, Ancoli-Israel S, Edinger JD, Krystal AD, Lichstein KL, Morin CM. The Consensus Sleep Diary: standardizing prospective sleep self-monitoring. Sleep. 2012;35(2):287-302. DOI 10.5665/sleep.1642. https://doi.org/10.5665/sleep.1642
  18. Morin CM, Belleville G, Belanger L, Ivers H. The Insomnia Severity Index: psychometric indicators to detect insomnia cases and evaluate treatment response. Sleep. 2011;34(5):601-608. PMID 21532953. https://pubmed.ncbi.nlm.nih.gov/21532953/
  19. Validation of the PROMIS Sleep Disturbance and Sleep-Related Impairment item banks. Sleep Medicine. 2022. PMID 35091171. https://pubmed.ncbi.nlm.nih.gov/35091171/ (verify authors, journal, volume, and pages before publication)
  20. Chinoy ED, Cuellar JA, Huwa KE, et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. 2021;44(5):zsaa291. DOI 10.1093/sleep/zsaa291. https://doi.org/10.1093/sleep/zsaa291
  21. Lee T, Cho C, et al. Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis. Journal of Clinical Sleep Medicine. 2025;21(3):573-582. DOI 10.5664/jcsm.11460. https://doi.org/10.5664/jcsm.11460 (verify authors before publication)
  22. Schyvens AM, Peters B, Van Oost N, et al. A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. SLEEP Advances. 2025;6(2):zpaf021. https://doi.org/10.1093/sleepadvances/zpaf021 (verify DOI, authors, and pages before publication)
  23. Kainec KA, Caccavaro J, Barnes M, Hoff C, Berlin A, Spencer RMC. Evaluating accuracy in five commercial sleep-tracking devices compared to research-grade actigraphy and polysomnography. Sensors. 2024;24(2):635. DOI 10.3390/s24020635. https://doi.org/10.3390/s24020635 (verify authors before publication)

Sign up for the Centralive Newsletter: https://newsletter.centralive.health/signup