Wednesday, October 7, 2026

DFA a1: From Lab to Consumer Exercise Prescription and Monitoring, ESSR 10/2026

It's been an interesting 6 years since I first published my "case report" proposing DFA a1 as a surrogate for the established endurance exercise thresholds. Since that time, the concept has gathered momentum with multiple publications by various groups. I've been fortunate to have been able to team up with Thomas Gronwald (originally) and, over the past few years, with Juan Murias and his (former) PhD student Pablo Fleitas-Paniagua. Along the way, we have studied some novel topics, including a revision of the first HRV threshold, effects of ramp slope on HRVTs, obtaining the HRVT2/MMSS via a submaximal ramp, demonstrating the potential of a1 as a measure of durability, validating Fatmaxxer and HRVTs under hypoxic conditions. Given the numerous positive studies demonstrating the utility of DFA a1 as a surrogate for established exercise thresholds and durability, we wrote an article advocating for widespread adoption of this measure in appropriate demographic groups. It is officially out today, in Exercise and Sports Science Reviews. Let's hope the folks at Garmin, Apple and Google take note, although we now have apps that will report DFA a1 in real time for each brand of device.

 

Text as follows: 

Blog Index.......

Thursday, September 24, 2026

Does Hilo Core Track Blood Pressure Accurately Across the Day?

As a physician concerned about blood pressure and a wearable enthusiast, I wanted to share some observations on the PPG device called Hilo Core. It may represent the most officially studied and "validated" PPG wrist device available (see PMID  33675592, 37016925, 38997475, and 39927495). After my previous post dealing with the considerable hurdles to overcome for this type of tech to be accurate, I was quite interested to see how it did. The main issue, though is not having "ondemand" BP capability. After looking over my data, I wonder if this is by design to make this type of comparison difficult. However, we can outsmart the Hilo app and do a simultaneous cuff vs wrist PPG relatively easily.

Here is how I did it: The Hilo band will only measure "BP" when you are stationary. This rules out anything while exercising or even perhaps if the HR is high. Next, the app segments the readings in 2 hour blocks:

 

Therefore, if we have the situation above, when 6 PM rolls around and no reading is done, we have an opportunity to do a simultaneous cuff reading. What seems to work is opening the app, and doing a manual cuff measurement:

 

Usually (if you are motionless), the wrist PPG will also do a measurement within a minute or two. If the 6 PM field is already populated, it will not show you the separate readings, so you will have to wait for the 8 PM block. You can also achieve the same concept by just keeping the band off the wrist until you are ready to do a cuff reading.

I began by doing the recommended calibration with the supplied cuff while in a quiet, resting, early AM state, before starting my (hectic) day. My goal was to see if the daytime reading "agreed" with cuff measurements based on the company recommended calibration procedure. As I discussed in the prior post, this may not be the best way to go about calibration.

As the paired reading came in, I was disturbed at how disparate the numbers were. The daytime values were after exercise, food, or just done randomly. About 10 days into this, I decided to recalibrate but do it during the day, which is when my measurements occurred. The results are as follows:

 

Calibration timeMeasurenCuff meanHilo meanMean differencePaired tp
DaytimeSystolic BP (mmHg)21124.7123.4−1.3−0.770.450

Diastolic BP (mmHg)2174.272.7−1.5−1.360.190

Heart rate (bpm)2163.463.5+0.10.400.691

Morning

Systolic BP (mmHg) 

19126.8115.3−11.5−5.65<0.001

Diastolic BP (mmHg)1974.170.5−3.6−3.730.0015

Heart rate (bpm)1965.465.1−0.3−1.190.250

 Systolic Bias:

 

  • There is a large bias with the morning calibration, which is not present with the daytime calibration group.

Diastolic Bias:

  • Bias is better than systolic, and daytime calibration is better than morning. 

Correlation:

 

 

  • Systolic regression is poor for morning calibration and poor for daytime calibration

  • Diastolic regression is good for both calibrations.   

 

Heart Rate:

 

 

  • Excellent agreement and correlation with HR cuff vs PPG

 

Key findings:

1. Systolic blood pressure showed a major calibration dependent difference

  • With the daytime calibration, Hilo systolic pressure averaged only 1.3 mmHg lower than the cuff.
  • With the morning calibration, Hilo averaged 11.5 mmHg lower than the cuff.
  • This approximately 10.2-mmHg difference in measurement bias between the two calibration periods was statistically significant:
  • The 95% confidence interval for the difference was approximately 4.8 to 15.6 mmHg.
  • This was the clearest finding in the dataset.

2. Daytime calibration largely removed the average systolic bias

  • Mean Hilo−cuff systolic difference: −1.3 mmHg

  • Paired comparison: p = 0.45

  • Therefore, there was no evidence of a systematic average difference between Hilo and the cuff during this period.

        During the morning-calibrated period:

  • Mean difference: −11.5 mmHg

  • Paired comparison: p < 0.001

Hilo therefore substantially underestimated systolic pressure during the morning-calibration period.

3. Better average agreement did not mean good agreement for individual readings

  • Although daytime calibration almost eliminated the average systolic bias, the Bland–Altman limits of agreement remained relatively wide:
  • Daytime calibration: −16.9 to +14.2 mmHg
  • For the morning calibration they were even wider:
  • −29.0 to +5.9 mmHg
  • This means that individual Hilo systolic measurements could still differ considerably from the cuff even when the average bias was close to zero.

In other words, calibration improved accuracy of the average but did not make individual Hilo and cuff readings interchangeable.

4. Systolic correlation remained poor under both calibration conditions

  • The correlation between Hilo and cuff systolic measurements was:
  • Daytime calibration: r = 0.29

  • Morning calibration: r = −0.19

  • Neither represents strong tracking of individual systolic changes.
  • Importantly, the difference between these two correlations was not statistically significant.

Thus, recalibration primarily improved the absolute bias, rather than producing a dramatic improvement in measurement-to-measurement correlation.

5. Concordance remained weak for systolic pressure

Lin's concordance correlation coefficient, which assesses both correlation and agreement, was:

  • Daytime calibration: CCC = 0.26

  • Morning calibration: CCC = −0.05

The daytime value represents an improvement, but agreement remained weak.

This again indicates that a good average value does not necessarily mean Hilo reliably tracks each individual cuff measurement.

6. Morning calibration showed evidence of proportional systolic bias

  • The Bland–Altman analysis also examined whether error changed depending on the actual blood pressure.
  • With morning calibration, there was significant proportional bias:
  • p = 0.022
  • This means the Hilo−cuff difference was not simply a constant offset; the amount of disagreement changed across the BP range.
  • With daytime calibration, proportional bias was weaker and narrowly missed conventional statistical significance:
  • p = 0.060
  • This suggests that daytime recalibration may have reduced both the fixed bias and some of the pressure-dependent error.

7. Diastolic pressure was less affected by calibration timing

  • The average Hilo−cuff diastolic difference was:
  • Daytime calibration: −1.5 mmHg

  • Morning calibration: −3.6 mmHg

  • The difference between calibration periods was approximately 2.1 mmHg, but this was not statistically significant:
  • p = 0.155

Therefore, the strong calibration-dependent effect seen for systolic pressure was not clearly present for diastolic pressure.

8. Diastolic correlation was considerably better than systolic correlation

  • Cuff and Hilo diastolic values correlated moderately well:
  • Daytime calibration: r = 0.60

  • Morning calibration: r = 0.63

  • This was substantially better than the systolic relationship.

However, morning-calibrated diastolic measurements still showed significant proportional bias, illustrating why correlation alone does not establish agreement.

9. Heart-rate agreement was excellent under both conditions

  • Heart rate served as an informative comparison because both devices measured HR at the same observations.
  • The Hilo−cuff HR difference was only:
  • Daytime calibration: +0.14 bpm

  • Morning calibration: −0.32 bpm

  • Correlations were extremely high:
  • Daytime: r = 0.94

  • Morning: r = 0.96

  • Lin's concordance coefficients were similarly strong:
  • Daytime: 0.94

  • Morning: 0.96

This suggests that the poor systolic agreement was not simply caused by gross temporal misalignment between measurements.

10. The data are consistent with a calibration-state effect

The original Hilo calibration was performed early in the morning, while many subsequent comparisons were performed during the daytime and around exercise.

After recalibration under daytime conditions, the average systolic error improved from approximately:

−11.5 mmHg to −1.3 mmHg

This raises the possibility that the relationship Hilo uses to estimate blood pressure may depend partly on the physiological state present at calibration.

Potential contributors could include changes in:

  • vascular tone

  • sympathetic activity

  • arterial stiffness

  • peripheral vasoconstriction or vasodilation

  • temperature

  • posture

  • recent exercise

A cuffless system based on pulse-wave characteristics may therefore perform best near the physiological conditions under which it was calibrated.

11. What does this mean for someone trying to determine whether they have mild hypertension?

This may be the most practically important implication of the experiment.

Under the current 2025 AHA/ACC blood-pressure classification, adult BP is categorized as:

  • Normal: systolic <120 and diastolic <80 mmHg

  • Elevated: systolic 120–129 and diastolic <80 mmHg

  • Stage 1 hypertension: systolic 130–139 or diastolic 80–89 mmHg

  • Stage 2 hypertension: systolic ≥140 or diastolic ≥90 mmHg

What is sometimes casually described as “mild hypertension” therefore generally corresponds to the Stage 1 range beginning at 130/80 mmHg. The guideline classification is based on averages of multiple careful measurements rather than a single isolated reading.

This is precisely the range in which measurement error can become clinically important.

For example, consider someone whose true cuff systolic pressure averages 132 mmHg, just inside the Stage 1 hypertension range.

With the −11.5-mmHg average systolic bias observed during the morning-calibrated portion of this experiment, a pressure of 132 mmHg could, on average, be represented by Hilo as roughly 120–121 mmHg.

That would move the apparent result from Stage 1 hypertension into the “elevated” category.

This calculation is only an illustration—the individual Hilo error cannot be predicted simply by subtracting the average bias—but it demonstrates the potential magnitude of classification error.

The opposite problem is also possible because the Bland–Altman limits were wide. Some individual Hilo readings were substantially above the corresponding cuff value.

Therefore, depending on the particular reading, a person near a diagnostic threshold could potentially be classified in either a higher or lower BP category.

12. Daytime calibration improves this problem—but does not completely solve it

After daytime calibration, the mean systolic bias was only −1.3 mmHg.

At the population-average level, that looks excellent. A cuff pressure of 132 mmHg would correspond to approximately 131 mmHg after applying the observed mean difference, leaving the measurement in the same Stage 1 category.

But the Bland–Altman limits remained approximately:

−16.9 to +14.2 mmHg

That range is considerably larger than the 10-mmHg width of the entire Stage 1 systolic category from 130 through 139 mmHg.

Consequently, a small average bias does not establish that an individual Hilo reading can reliably answer the question:

“Is my true blood pressure below or above 130 mmHg?”

The same issue applies to the 80-mmHg diastolic threshold.

For someone whose actual diastolic pressure sits near 80–85 mmHg, an error of only several mmHg could move the displayed result between normal/elevated and Stage 1 hypertension.

13. Classification is different from trend monitoring

This highlights an important contrast between two potential uses of a wearable BP device.

One use is:

“Is my blood pressure generally going up or down over weeks or months?”

A device might potentially provide useful longitudinal information even with some individual measurement error.

A much more demanding use is:

“Is my actual average BP 126, 132, or 142 mmHg, and therefore which hypertension category am I in?”

That requires sufficient absolute accuracy around clinically important thresholds.

The results of this N-of-1 experiment suggest caution about using Hilo alone for the second question, particularly when the user's BP is close to 130/80 mmHg.

14. This concern is consistent with current AHA guidance on cuffless devices

The 2025 AHA/ACC hypertension guideline specifically states that reliance on cuffless devices, including smartwatches, for accurate BP measurements should be avoided until they demonstrate greater precision and reliability.

The AHA's subsequent scientific statement on cuffless BP devices notes that these systems usually estimate BP indirectly, often relative to a calibration measurement, and that accuracy can be affected by factors including calibration, motion, posture, and real-world physiological conditions.

For home BP assessment, the AHA currently recommends a validated automatic upper-arm cuff device. It recommends taking multiple measurements under standardized resting conditions and recording results over time.

That recommendation is particularly relevant for someone wondering whether their average BP is just below or just above the Stage 1 hypertension threshold.

Practical implication for someone near 130/80 mmHg

If someone's Hilo readings consistently average, for example, 125/76 mmHg, 132/81 mmHg, or another value close to a diagnostic boundary, this experiment suggests that the wearable value alone may not be sufficiently precise to determine whether that individual truly has Stage 1 hypertension.

A sensible approach would be to use a validated upper-arm cuff as the reference measurement, under standardized conditions:

  • rest quietly for at least five minutes

  • avoid exercise, caffeine and smoking for at least 30 minutes beforehand

  • support the arm at heart level

  • take two readings approximately one minute apart

  • repeat measurements across multiple days

For determining whether a person meets criteria for hypertension, reference BP should be measured under standardized resting conditions. However, evaluating the real-world validity of a cuffless wearable requires a different approach. Comparisons should also include physiologically challenging conditions—such as daytime activity, changes in posture, and post-exercise recovery—because these are precisely the situations in which vascular tone and pulse-wave characteristics may depart most from the calibration state. 

These procedures are consistent with current AHA home-monitoring recommendations.

The 2025 guideline also emphasizes that BP classification is based on an average of at least two careful readings obtained on at least two occasions, rather than an isolated measurement.

For someone close to 130/80, the average of properly obtained cuff measurements is therefore considerably more appropriate for answering “Do I meet the definition of hypertension?” than an isolated or cuffless wearable reading.

Overall conclusion

The most important observation was that Hilo systolic accuracy changed markedly after recalibration.

Morning calibration was associated with an average systolic underestimation of approximately 11.5 mmHg, whereas daytime calibration reduced that error to approximately 1.3 mmHg.

However, even after recalibration, Bland–Altman limits remained wide and correlation with individual cuff systolic measurements remained limited.

For someone whose main question is whether they have borderline or Stage 1 hypertension, these findings matter because the AHA threshold begins at 130/80 mmHg, while the observed individual disagreement between the Hilo and the cuff was often considerably larger than the few millimeters of mercury separating one BP category from another.

Thus, the data suggest that calibration may strongly influence Hilo's average systolic BP accuracy, while the remaining measurement-to-measurement variability limits confidence in using the wearable alone to determine whether an individual's true BP lies just above or below a hypertension threshold.

For that purpose, current AHA guidance continues to favor repeated measurements using a validated automatic upper-arm cuff.

 

Why should we be concerned about exercise and recovery BP dynamics?

An exaggerated blood pressure response during exercise has been associated with masked hypertension, future development of hypertension, and higher cardiovascular risk, even in people whose resting BP appears normal.

A 2023 meta-analysis found that people with masked hypertension were more than three times as likely to show an exaggerated BP response to exercise compared with truly normotensive individuals (OR 3.33), suggesting that exercise BP can expose hypertension that is not apparent at rest (PMID: 36980313).

Longitudinal studies support this relationship. In 3,742 initially normotensive men, a peak exercise systolic BP above approximately 181 mmHg predicted a higher risk of developing hypertension over about five years (PMID: 25824452). An earlier study of 5,386 normotensive men likewise found that an exaggerated exercise BP response independently predicted future hypertension after accounting for resting BP and other risk factors (PMID: 9467632).

A systematic review including more than 35,000 normotensive participants concluded that exaggerated exercise BP is associated with both future hypertension and cardiovascular events (PMID: 28511070). Exercise SBP has also been associated with adverse vascular and metabolic characteristics, including greater arterial stiffness and poorer endothelial function (PMID: 34738988).

Importantly, some investigators argue that submaximal exercise BP may be particularly useful, because moderate workloads resemble the cardiovascular demands of normal daily activity more closely than maximal exercise. Practical guidance has therefore proposed measuring BP at a standardized moderate workload as a way to identify individuals whose hypertension may not be evident from resting measurements alone (PMID: 35270514). 

BP contextEstablished role
Standardized seated BPPrimary classification/diagnosis
Home BPDiagnosis/management
24-h daytime ABPMEstablished diagnosis + risk prediction
Submaximal exercise BPRisk marker; can uncover masked HTN
Peak exercise BPRisk marker, but thresholds less standardized
Recovery/post-exercise BPEmerging/established prognostic marker, not diagnostic criterion
BP during unrestricted activityInteresting physiologically, less standardized

 

Hilo PPG BP, the bottom line:

Cons:

  • Poor agreement and correlation to daytime BP when calibrated per recommendation (in AM)
  • Minimal bias, but still a large limit of agreement and weak correlation when calibrated for daytime.
  • Can easily miss mild HTN
  • Only measures at rest with no motion
  • Can't use pre/post/during exercise 
  • Difficult to do on-demand readings.
  • AHA recommends avoidance of PPG BP usage
  • Cost and subscription model

Pros:

  • Minimal footprint and weight
  • Does not require inflation (especially important at night)  
  • If calibrated properly, overall average BP does match with cuff 

 

Relevance to Smartwatch BP:

The major smartwatch companies are now moving aggressively into blood-pressure monitoring, but their approaches differ substantially.

Google's Pixel Watch 5 has begun rolling out a new Blood Pressure Trends feature as part of Google Health's “Health Guardian” system. Rather than displaying conventional systolic and diastolic numbers, the watch analyzes physiological signals continuously and looks for changes in blood-pressure patterns over time. Google describes this as a trend-monitoring feature rather than a replacement for cuff measurements. The feature is also being extended to Pixel Watch 3 and 4.

Apple Watch takes a similar approach. Apple now offers hypertension notifications on supported watches. The optical heart sensor analyzes vascular responses over approximately 30-day periods and can notify a user when it detects a pattern consistent with chronic hypertension. Importantly, Apple does not provide an estimated 128/82-style BP reading. If a hypertension notification occurs, Apple recommends confirming the finding with a conventional BP cuff over seven days.

Samsung takes a different approach. Its Galaxy Watch blood-pressure feature is now available to U.S. users on compatible watches and provides estimated systolic and diastolic BP values. Like Hilo, however, it requires calibration against an upper-arm cuff and requires recalibration every 28 days. Samsung explicitly states that the feature is not intended to diagnose disease. Samsung has also announced passive blood-pressure trend monitoring for later in 2026.

Accuracy issues become particularly important near diagnostic boundaries such as 130/80 mmHg, where an error of only several millimeters of mercury can change BP classification.

For this reason, claims that a smartwatch can detect hypertension, monitor blood-pressure trends, or estimate systolic and diastolic pressure should not automatically be interpreted as evidence that it can replace a validated upper-arm cuff.

The critical question is not simply whether the device performs well in a controlled resting validation study. It is whether it remains accurate across the conditions in which people actually intend to use it—morning, daytime activity, sleep, emotional stress, exercise and post-exercise recovery—and whether it can correctly identify people close to clinically important BP thresholds.

Until independent studies demonstrate that level of real-world performance, smartwatch BP and hypertension features are best viewed as screening or supplementary tools rather than definitive measurements of blood pressure or proof that hypertension is present or absent.

Blog Index

Tuesday, September 8, 2026

Watchletic, DFA a1 for the Apple Watch and Wear OS

Updated 9/11/26 - eval of version 3.2 beta

Every so often a new sports monitoring app catches me by surprise. As a Pixel Watch user, I have a weekly ChatGPT search scheduled for any developments in HRV or Bluetooth HRM for Wear OS. A few weeks ago I was pleasantly surprised by a chat notice that a "new-ish" app now supported Bluetooth HRV and DFA a1 for Wear OS (and the Apple watch). 

 

 

 

The app is called Watchletic, and in this post we will take a quick peek at its ability to track DFA a1 during a sample indoor trainer session. Since the app is still in development and the DFA a1 implementation is still being fine tuned, no formal stats will be presented. However, in a future post, a more formal evaluation will be done. As usual, this is N=1 so YMMV.

The app was installed on my PW5 watch with no issues and pairing with a Polar H10 was quick. The screen interface is also straightforward to use to create custom fields on your watch:

 


Here is a screenshot of the entire session from my Fold 8:

Here is a different session showing just power and DFA a1:

 

The lines are color coded with stats displayed where you put the cursor (see the arrow above) The recorded file is downloadable as a .fit file. That .fit file can be uploaded or analyzed as well, here is the file data as interpreted by Runalyze:

 


  • It should be noted that the Runalyze file DFA a1 values/plot is derived from the RR output of the H10 (from the .fit file). I used that same .fit file for Kubios vs Watchletic a1 for the comparison.

The following 2 hour session was comprised of a warm-up, a submax ramp (up to my first threshold), and 5 sets of blood flow restriction (2 min on - 2 min off with compression at 75% AOP).

Since we have Runalyze on board, looking at the O2 sat drops through the BFR is always fun:

 

(Yeah, the first set had minimal O₂ desaturation.) 

How did DFA a1 fare between the RRs analyzed with Kubios vs Watchletic?

Here is the entire 2 hours. The max artifact was 1% or less, so we don't have to be too concerned about the "better" Kubios automatic correction method except in one spot, as we will discuss below.

 

A couple of comments before we zoom in to the ramp and BFR intervals:

  • The output data is now in a CSV, eliminating previously present timing issues. Many thanks to the developer for doing this.
  • The app does use Kubios detrending methodology and threshold artifact correction. My default for Kubios is "automatic" correction, which is more advanced. This is important as the alignment between Kubios and Watchletic deviated in the ramp.

Zooming in to the ramp/intervals:

 

    Artifact correction threshold enabled in Kubios

     

  • The green circle shows the deviation. Why does this happen? It turns out to be purely from different artifact correction methods. If I switch Kubios over to "threshold" methodology, that transient a1 rise shows up, just as in Watchletic.This is important to remember - if you have many artifacts, most software will have a DFA a1 bias upward.

Blood flow restriction sets:

 

  • This is the most encouraging image of the bunch. Why? It shows virtually identical dynamic range of a1 from high to low on the BFR intervals. It's been my experience that apps that don't do well vs Kubios have little issue in the 0.8 to 1.5 range but can't reach the low nadirs of the anticorrelated ranges below 0.5.

What else can the app do, HRV wise?

It also has the option of reporting RMSSD for those interested in resting HRV. It's still using 2 minute windows but can display the running average:

 

  • The RMSSD (and artifact %) as a data field will show on the watch display, making it useful for resting and recovery usage:

 

  • The "Next" circle is a button press for easy interval demarcation and interval timer reset

How are the comparison stats?

The following stats are from 2 sessions: the Ramp/BFR/5 minutes in high zone 2 and another session in zone 1, with a 5 minute mid-zone 2 interval.

Ramp/BFR (2 hrs):

Regression:

 

Bland Altman with minimal proportional bias (at high a1):

 

Zone 1 session with a 5 min zone 2 (60 min):

Regression:

 

Bland Altman:

 

  • Both sets of raw data are not normally distributed, so the regression is "unofficial" (for you "real" science people out there), but I think you can see that the correlation is very good to excellent.
  • The Bland Altman plots also show good agreement with a tiny bit of bias, possibly due to the artifact correction issues mentioned above. Both sessions were similar in this regard. 

Summary:

  • Apple Watch and Wear OS users now have an option for DFA a1 display and RR recording  while exercising. This is a very reasonable option to both display DFA a1 in real time, record RRs for later upload and show the session graphics on your phone afterward.
  • For those interested in RMSSD, artifact and cycling power that is also available in the smartphone and watch display. 
  • So far the "agreement" with Kubios HRV is excellent and at least in my case, similar to Fatmaxxer.
  • Bottom line - this is an exciting development in mobile HRV and DFA a1. Use cases include times when you can't take your phone but have a Wear OS or Apple Watch (trail run, distance run). Additionally, if one wanted to monitor a1 and RMSSD (and artifacts) inexpensively, get this app and a used Pixel watch 3-4 or used Apple Watch, and you are good to go!

All tested DFA a1 apps

 ....Blog Index....

Thursday, August 20, 2026

Hilo, Oura, and Cuffless Blood Pressure: Promise, Physiology, and Pitfalls

The PPG, Blood Pressure Paradox: Why Hilo and Oura May See the Pulse but Not the Pressure:

With the recent wave of wearable devices claiming some form of blood-pressure tracking, from Oura’s nighttime BP dipping to Whoop, Galaxy Watch, and now even the Pixel Watch 4/5, it seems like a good time for a reality check: what are these devices actually measuring, what do their BP-related numbers really mean, and how accurate are they?. To that end, first we will discuss the theoretical background of optical BP and then mention a more "dedicated" unit, the Hilo Core wristband vs the Oura ring, which I’ve used for the past few weeks.

What does a PPG blood-pressure device actually measure?

Photoplethysmography (PPG) has become one of the most widely used sensing technologies in consumer wearables. By lighting the skin with an LED and measuring changes in reflected or transmitted light, PPG detects the pulsatile changes in blood volume that accompany each heartbeat. This signal is well suited to estimating heart rate and, under favorable conditions, beat-to-beat timing for heart-rate variability. More recently, however, PPG has been extended to a much more ambitious target: blood pressure estimation without an inflatable cuff. This development has considerable appeal. Conventional ambulatory blood-pressure monitoring requires repeated cuff inflation, which can be uncomfortable, disruptive during sleep, and limited to intermittent measurements. I can attest to this personally. Several months ago I took part in a Pixel watch ambulatory BP study and wore a cuff for 24 hours. Every time it inflated, it woke me from sleep and I ended up just removing it before it drove me nuts. Therefore, a small wristband, watch, or ring capable of estimating blood pressure passively could potentially provide many measurements across days or weeks with little burden on the wearer. Devices such as Aktiia/Hilo and newer PPG-based features from Oura illustrate how rapidly this field is moving.

The important physiological distinction is that PPG does not measure pressure directly. A conventional cuff estimates arterial pressure by mechanically compressing an artery and observing the pressure at which arterial flow or oscillations change. A PPG sensor instead records an optical waveform generated by pulsatile changes in peripheral blood volume. Blood pressure influences that waveform, but it is only one of many factors that do so.

The shape of a PPG pulse contains substantial cardiovascular information. Features such as the systolic upstroke, systolic peak, dicrotic notch, reflected or diastolic wave, pulse width, and timing between waveform landmarks can change as arterial pressure changes. These relationships provide the physiological foundation for cuffless BP algorithms. Modern systems may combine many such waveform characteristics, along with heart rate and other sensor information, using statistical or machine-learning models to infer blood pressure or blood-pressure-related patterns.

The difficulty is that these waveform characteristics are not uniquely determined by blood pressure. Peripheral vascular tone, arterial stiffness, pulse-wave velocity, wave reflection, stroke volume, heart rate, skin temperature, local perfusion, body position, autonomic activity, and even the pressure of the sensor against the skin can modify the PPG waveform. Thus, a change in pulse shape does not necessarily mean that blood pressure changed by a corresponding amount.

Conceptually, the PPG signal might be represented as:

PPG waveform = a function of: (blood pressure, vascular tone, arterial stiffness, wave reflection, stroke volume, heart rate, temperature, perfusion, sensor conditions, …)

A cuffless BP algorithm attempts to solve the inverse problem:

blood pressure ≈ summated function⁻¹(PPG waveform)

That inverse relationship is considerably more difficult. If several physiological variables can produce similar changes in the optical pulse, there may be no unique blood-pressure value associated with a particular waveform. The problem becomes especially important when the physiological state changes substantially, for example during sleep, exercise, medication treatment, heat exposure, or sympathetic activation. This distinction helps explain an apparent paradox in the cuffless-BP literature. A PPG device can demonstrate superb agreement with a conventional cuff during controlled resting measurements yet perform much less well when asked to follow blood pressure changes over time. Calibration can place the estimated pressure close to the true value at baseline, but good baseline agreement does not necessarily demonstrate that the algorithm has the correct "dynamic response" or "gain" when blood pressure subsequently rises or falls.

Nocturnal blood-pressure dipping provides a particularly useful example. During sleep, blood pressure normally declines, but so do sympathetic activity, peripheral vascular resistance, heart rate, skin temperature, and other determinants of the peripheral pulse waveform. A PPG algorithm therefore must distinguish waveform changes caused specifically by falling arterial pressure from waveform changes caused by the broader transition from wakefulness to sleep.

This issue is not confined to devices that report systolic and diastolic pressure in mmHg. It also applies to algorithms that use PPG to classify people as normal dippers or non-dippers. Such systems may avoid the difficult task of reconstructing absolute blood pressure, but they still depend on a relationship between nocturnal pulse morphology and the underlying blood-pressure response.

The central question for PPG-based blood-pressure technology is therefore not whether the optical pulse contains information related to blood pressure, it clearly does. The more difficult question is whether an algorithm can reliably separate the effect of blood pressure from the many other physiological processes that alter the PPG waveform, particularly when the cardiovascular state changes over time.

Hardware/Software Considerations:

Sampling requirements differ substantially between ECG and PPG because the important information in the two signals is different. In an ECG, heart-rate variability depends primarily on accurately locating the R peak, a sharp and readily identifiable electrical fiducial point. At a sampling rate of 250 Hz, samples are separated by 4 ms; at 500 Hz by 2 ms; and at 1000 Hz by 1 ms. Studies using ECG signals down-sampled from 1000 Hz have generally found excellent agreement in HRV measures at 250–500 Hz, while lower sampling rates progressively increase timing quantization error. Interpolation around the R peak can improve timing considerably and permit useful HRV analysis even at lower sampling rates. 

PPG presents a different problem. Beat timing can be obtained from a single reproducible point on the pulse, such as the foot, maximum upslope, or systolic peak and reasonable pulse-rate variability can therefore be obtained at surprisingly low sampling rates. One study found that, with an interpolated upslope fiducial point, PPG could be reduced to 50 Hz without important changes in conventional variability indices. This does not, however, mean that 50-Hz PPG adequately reproduces pulse morphology.

For blood-pressure estimation, the algorithm may depend on the shape of the entire pulse: the systolic upstroke, systolic peak, dicrotic notch, reflected wave, and the timing and amplitude relationships among these features. At 50 Hz there is only one sample every 20 ms, and at 100 Hz one every 10 ms. Thus, a relatively small feature such as the dicrotic notch may be represented by only a few samples, and its apparent position can shift substantially depending on where it falls relative to the sampling grid. Interpolation can improve estimates of fiducial timing, but it cannot restore morphological information that was never acquired, particularly in the presence of noise or waveform distortion. Work specifically examining PPG fiducial localization confirms that sampling rate, filtering, interpolation, and signal quality all influence the localization of waveform landmarks used for BP estimation.

Consequently, 100 Hz is quite reasonable for gross PPG morphology and is the documented sampling rate of the validated Hilo system, but higher rates such as 200–250 Hz provide substantially denser representation of the pulse when precise timing of the dicrotic notch, reflected wave, or other morphological features is important. This is different from ECG R-wave detection, where a single sharp fiducial point can often be localized very accurately even from a moderately sampled signal. A high sampling rate therefore does not make PPG-derived blood pressure intrinsically accurate, but inadequate sampling can add yet another source of uncertainty to an already indirect physiological measurement. That said, Oura's historically reported 250-Hz nocturnal PPG could provide a prettier and more precisely sampled pulse waveform than Hilo’s 100-Hz signal, yet it still cannot solve the fundamental BP problem. Sampling improves measurement of the PPG waveform; it does not make PPG morphology uniquely determined by arterial pressure.

Calibration: Why Good Agreement Does Not Necessarily Mean Good BP Tracking

Most PPG-based blood-pressure systems face an immediate problem: the optical waveform has no intrinsic scale in mmHg. A PPG sensor measures changes in reflected light associated with pulsatile blood volume, producing a waveform in arbitrary optical units. Some method is therefore needed to connect the optical signal to an actual arterial pressure. The simplest solution is calibration against a conventional cuff. This distinction is particularly clear in the published Hilo/Aktiia validation studies. The bracelet first analyzed the wrist PPG waveform to produce uncalibrated systolic and diastolic BP estimates. During initialization, the first paired cuff measurements were then used to calculate a subject-specific SBP offset and DBP offset, converting the optical estimates into mmHg (PMID: 33675592). Conceptually, this can be represented as:

Reported BP = PPG-derived BP estimate + individual calibration offset

Suppose the optical algorithm estimates an individual's resting systolic pressure as 112 arbitrary BP units, while the calibration cuff measures 124 mmHg. An offset of +12 mmHg can make the bracelet report:

112 + 12 = 124 mmHg

The device now agrees perfectly with the cuff at calibration. But that agreement tells us surprisingly little about what will happen when the person's BP changes. Imagine that several hours later the true systolic pressure falls by 20 mmHg:

Cuff: 124 → 104 mmHg

If the PPG morphology changes by an amount that the optical algorithm interprets as only a 6-mmHg fall, the device will calculate:

112 → 106 optical units

and, after applying the same +12-mmHg calibration:

124 → 118 mmHg

So…. The initial value is exactly correct, but the subsequent change is badly underestimated. The problem is not the calibration offset. The problem is that the dynamic sensitivity of the PPG model to changing BP is too small. An offset calibration primarily solves the problem: Where should the BP estimate sit on the mmHg scale. It does not necessarily solve the problem: How far should the BP estimate move when actual pressure changes?

If the correct gain is 1.0 but the PPG algorithm effectively behaves as though it were 0.3, a true 20-mmHg change may appear as only about a 6-mmHg change. No adjustment of the baseline can correct that. This is especially relevant because Aktiia's later 24-hour study describes these two functions explicitly as separate. The investigators state that initialization generates an offset, establishing the absolute reference around which subsequent measurements vary, whereas the device's ability to track BP trends is determined independently by the optical algorithm. Reinitialization changes the baseline but does not change the trend-tracking mechanism (PMID: 39927495).

Population accuracy is not the same as within-person tracking

This also illustrates an important limitation of conventional device-validation statistics. Suppose 100 people undergo seated testing. Each device is individually calibrated near the subject's resting BP. Across the population there may be people with pressures ranging from 100 to 180 mmHg. If calibration places each person's optical estimate near his or her cuff pressure, the device can show:

  • ·       very small mean bias,
  • good correlation across subjects,
  • acceptable average error,
  • an impressive Bland–Altman plot.

The original Aktiia validation study, for example, reported average differences of approximately 0.46 mmHg for SBP and 0.39 mmHg for DBP during standardized seated measurements after subject-specific calibration (PMID: 33675592). Those are encouraging results, but they primarily demonstrate that the system can provide reasonable calibrated absolute BP estimates under the tested conditions. They do not automatically demonstrate that, within a particular individual:

ΔBP from PPG ≈ ΔBP measured by the reference method.

This is a fundamentally different validation question. A device could therefore show excellent cross-sectional agreement because subjects with high BP generally receive high estimates and subjects with low BP receive low estimates, while still substantially underestimating the magnitude of BP changes occurring within each person.

Why estimating an individual slope is much harder

In principle, calibration could determine both the intercept and the slope of the true BP vs PPG waveform value. But this requires reference measurements obtained across a meaningful range of BP. If cuff calibration measurements are: 121, 123, 120, 124 mmHg they provide excellent information about the person's resting baseline but very little information about how the PPG should behave at: 100 or 150 mmHg.

To estimate an individual PPG-to-BP slope reliably, paired PPG and reference measurements would ideally span substantially different pressure states.  That might require controlled perturbations such as:

  • ·       changes in posture,
  • exercise or recovery,
  • pharmacological BP changes,
  • cold or heat exposure,
  • sustained handgrip,
  • lower-body pressure manipulation,
  • or naturally occurring day-night changes.

But this introduces the next problem: these interventions alter vascular physiology as well as BP. The relationship between PPG morphology and pressure may itself change between conditions. Consequently, even a personalized linear slope may not be sufficient. The true relationship may be state dependent: BP = function of (PPG, vascular tone, temperature, posture, arterial stiffness, autonomic state, …) rather than a single fixed calibration equation.

This concern is supported by the independent evaluation of Aktiia by Tan and colleagues, in which antihypertensive treatment produced an approximately 19.7/11.5 mmHg reduction by cuff-based home BP monitoring, while Aktiia detected only about 1.0/0.8 mmHg of change (PMID: 37016925). Although this medication-change subgroup was very small and requires replication, the result is important because it illustrates precisely the distinction between obtaining a reasonable calibrated BP value and accurately tracking a substantial within-person BP change.

The same issue appears in nocturnal measurements. In a later 24-hour comparison, conventional ambulatory monitoring showed a substantially larger nighttime systolic decline than Aktiia, despite relatively close agreement in daytime mean pressure (PMID: 39927495). An independent ambulatory comparison similarly found marked underestimation of normal nocturnal BP dipping (PMID: 37016925). These findings raise the possibility that the optical system can track the *direction* of some BP changes while compressing their magnitude.

Thus, the crucial question for a calibrated PPG blood-pressure system is not simply: “Does its BP value agree with the cuff after calibration?” It is: does the optical signal correctly track both the direction and the magnitude of subsequent changes in blood pressure? The two are not equivalent—and distinguishing them is essential when interpreting validation studies of cuffless BP technology.

Oura Nighttime Blood Pressure: A Different Approach, but the Same PPG Limitation

Oura has taken a somewhat different approach to PPG-based blood-pressure assessment than Hilo/Aktiia. Rather than attempting to provide systolic and diastolic pressure in mmHg, Oura's Nighttime BP feature analyzes nocturnal PPG patterns and estimates the degree to which blood pressure appears to fall during sleep. The distinction is significant. Oura is not claiming that the ring directly measured, for example: Daytime SBP = 126 mmHg → nighttime SBP = 108 mmHg. Instead, it classifies the nighttime pattern into categories such as: 

  • ·       typical dipping: 10–20%
  • reduced dipping: <10%
  • pronounced dipping: >20%
  • rising: no fall or an increase overnight.

The displayed pattern is based on approximately 30 nights of data, rather than a single night's recording. Oura explicitly states that nighttime BP is a wellness feature, does not display BP in mmHg, and should not replace conventional blood-pressure monitoring. This represents a more modest target than reconstruction of absolute BP. However, it does not eliminate the fundamental physiological problem inherent in PPG-derived blood pressure: Oura is still inferring BP from a waveform that is not determined solely by BP!

As noted, the optical pulse recorded at the finger is affected by blood pressure, but also by many other variables (vascular tone, arterial stiffness, wave reflection, stroke volume, heart rate, temperature, peripheral perfusion, autonomic activity, sensor conditions, …) All of these factors can change during the transition from wakefulness to sleep. Therefore, if Oura observes a nocturnal change in pulse-wave morphology, that change cannot automatically be attributed to a corresponding change in arterial pressure. This is the same fundamental inverse problem encountered by other PPG-based BP systems. The American Heart Association's recent scientific statement on cuffless BP measurement emphasizes that PPG-derived pressure estimates depend on physiological relationships that may be altered by changes in vascular properties and measurement conditions and concludes that currently available cuffless devices have not yet been adequately validated across all typical use conditions (Cohen et al., 2026; PMID: 41376592).

What Oura's validation actually shows

 Oura reports that its science team compared Oura Ring 4 PPG data with 48-hour ambulatory BP monitoring in 134 participants. For discrimination between dippers and non-dippers, Oura reports:

  • ·         Sensitivity: 84%
  •  Specificity: 69%
  •  Area under the ROC curve: 0.87

These results are encouraging for a noninvasive screening signal. However, as of August 2026, this 134-participant Oura validation has not been identified as a PubMed-indexed peer-reviewed publication, and therefore no PMID is currently available. The performance figures are reported by Oura. The results also need to be interpreted according to what was actually tested. An AUC of 0.87 means that the PPG-derived features contain substantial information related to ABPM-defined dipping status. It does not mean that Oura measured nighttime blood pressure with 87% accuracy. Likewise, the reported 69% specificity means that, at the chosen classification threshold, approximately 31% of ABPM-defined non-dippers would not be correctly identified as non-dippers. That is meaningful uncertainty if the feature is interpreted at the level of an individual. More importantly, important methodological details remain unavailable or insufficiently described publicly, including:

  • ·         the exact PPG features entering the dipping model,
  •  the sampling rate used specifically for the Nighttime BP algorithm,
  •  signal filtering and pulse-rejection procedures,
  •  the proportion of nighttime data rejected for poor signal quality,
  • Whether sleep stages enter the model,
  •  the continuous relationship between estimated and ABPM-measured percentage dipping,
  •  Bland–Altman limits for percentage dipping,
  •  and performance in clinically important subgroups.

Until those data are published, the reported sensitivity, specificity, and AUC are best regarded as promising preliminary validation, rather than definitive evidence of quantitative BP-dipping accuracy.

Is the finger a better site for BP:

There are reasons why Oura could potentially perform better than wrist PPG. The finger generally provides a stronger peripheral pulse signal, and Oura has demonstrated very good performance for nocturnal heart rate and interbeat-interval measurement. But obtaining a high-quality waveform and interpreting its physiological meaning are two different problems. Higher sampling frequency, lower motion artifact and a clearer dicrotic notch can provide: a more accurate measurement of the PPG waveform without necessarily providing a more accurate measurement of blood pressure.

The Aktiia studies are particularly relevant because they demonstrate experimentally that PPG morphology can contain enough information to produce plausible BP estimates while still substantially underrepresenting nocturnal BP changes. That physiological limitation does not disappear when the PPG sensor is moved from the wrist to the finger.

Thirty night averaging: potentially an advantage, but also a different measurement

Oura's decision to average nighttime patterns over approximately 30 nights is potentially valuable. Dipping status itself is not perfectly reproducible from night to night. Burgos-Alonso and colleagues examined 225 high-cardiovascular-risk patients who underwent four 24-hour ABPM recordings over five months. Only about two-thirds of participants retained their initial systolic dipper/non-dipper classification across subsequent recordings, leading the authors to characterize individual dipping reproducibility as modest (Burgos-Alonso et al., 2021; PMID: 33591600). A second large study by McGowan and colleagues examined 512 untreated patients with two ABPM studies approximately 29 months apart. Binary dipper/non-dipper classification remained unchanged in 76% of participants, but agreement was relatively weak. In contrast, nocturnal dipping expressed as a continuous percentage was considerably more reproducible, supporting the concept that dipping may be better regarded as a continuous physiological variable rather than a rigid categorical state (McGowan et al., 2009; PMID: 19641455). Repeated passive PPG measurements, therefore, offer a genuine potential advantage: they can characterize the person's typical nocturnal physiology over many nights without repeatedly inflating a cuff and disturbing sleep. However, this also raises an important validation issue.

Oura's reported reference study used 48-hour ABPM, whereas the consumer feature characterizes patterns over approximately 30 nights. Those are not necessarily the same phenotype. A 30-night PPG average could conceivably be a more stable measure of nocturnal vascular physiology than a one- or two-night ABPM classification—but demonstrating that would require prospective outcome studies, not simply agreement with a short ABPM recording.

Oura's approach is conceptually more conservative than attempting to reconstruct systolic and diastolic BP in mmHg. A PPG-based classifier may indeed identify useful nocturnal cardiovascular patterns even when PPG cannot accurately reproduce absolute blood pressure. But the underlying physiological limitation remains: The nocturnal PPG waveform is affected by blood pressure, but it is also strongly affected by the same autonomic and vascular changes that accompany sleep. Oura may therefore be measuring a valuable surrogate of nocturnal BP physiology, rather than nocturnal BP itself. That distinction should remain central when interpreting the new Nighttime BP feature.

Some screen shots from my Hilo Core and Oura 4 ring from the same night:

Oura 4 

compared to the Hilo:



Both show some nocturnal dips, but as noted above, YMMV. 
I do like the form factor of the Hilo and hope to see improvements in calibration over time.

 

Conclusion:

  • PPG-based blood-pressure technology is attractive because it promises something conventional cuff measurement cannot easily provide: frequent, unobtrusive cardiovascular monitoring across daily life and sleep. The optical pulse clearly contains information related to arterial pressure, and increasingly sophisticated algorithms can extract timing, amplitude, and morphological features that correlate with BP and vascular state. 
  • The central limitation, however, is physiological rather than merely technical. A PPG sensor does not measure arterial pressure directly. It measures changes in peripheral blood volume, and the resulting waveform reflects the combined influence of blood pressure, vascular tone, arterial stiffness, pulse-wave reflection, stroke volume, heart rate, temperature, perfusion, autonomic activity, body position, and sensor conditions. Blood pressure is therefore one contributor to the waveform, not its sole determinant. That distinction has practical consequences. 
  • A device can agree closely with a cuff after calibration without necessarily tracking subsequent BP changes accurately. Calibration may establish the correct baseline or offset while leaving errors in the dynamic gain of the PPG-to-BP relationship. The Aktiia/Hilo literature illustrates this particularly well. Controlled seated measurements can show minimal average bias, while independent and ambulatory studies have demonstrated substantial attenuation of nocturnal BP dipping and, in limited data, large medication-induced changes in BP (PMID: 33675592; PMID: 37016925; PMID: 39927495). This should change how cuffless BP validation is interpreted. Low mean bias around the calibration condition is important, but it does not answer the more physiologically relevant question: Does the device correctly reproduce the magnitude of within-person BP changes across different physiological states? For many applications, including nocturnal dipping, medication response, BP variability, exercise recovery, and stress responses—that may be the more important test.
  • Nighttime monitoring is especially challenging because sleep alters many of the same physiological variables that determine PPG morphology. Sympathetic activity, vascular resistance, arterial wave reflection, skin temperature, peripheral perfusion, heart rate, and stroke volume all change during sleep. A nocturnal change in the optical pulse therefore cannot automatically be interpreted as a proportional change in arterial pressure. This limitation also applies to devices such as Oura that take a more conservative approach and infer BP-dipping patterns rather than reporting absolute SBP and DBP. Classification is an easier problem than reconstruction of pressure in mmHg, and a PPG-derived dipping phenotype may ultimately prove useful. But successful classification does not establish that the ring directly measured the magnitude of the BP decline. The algorithm may instead be detecting autonomic and vascular changes that are associated with normal dipping.
  • Signal quality and sampling rate matter, but they do not solve the underlying physiology. A higher sampling rate can improve identification of the systolic peak, dicrotic notch, reflected wave, and other pulse landmarks. Better sensor placement can reduce noise and motion artifact. More sophisticated machine learning can identify complex patterns that simple regression cannot. Yet none of these improvements make PPG morphology uniquely determined by blood pressure. Likewise, aggressive signal-quality filtering can create another potential source of bias. Wearables commonly reject periods containing motion, poor perfusion, atypical pulse morphology, or inadequate sensor contact. This improves technical signal quality but may also preferentially exclude unusual physiological states. A device may therefore provide very clean estimates during selected low-motion periods while incompletely representing the full variability of blood pressure across the day or night.
  • For future validation studies, the most informative analysis may therefore be less about absolute agreement after calibration and more about tracking change. Studies should deliberately expose participants to a broad range of pressures and physiological states. None of this means that PPG-based BP technology lacks value. On the contrary, PPG may ultimately prove extremely useful for long-term cardiovascular phenotyping, detecting trends, identifying individuals who warrant formal BP evaluation, and recognizing physiological patterns that conventional intermittent measurements miss. A multi-night optical phenotype may even prove more reproducible or prognostically informative than a single night of cuff-based monitoring.
  • “Can PPG predict blood pressure?” It clearly can to some degree. The more important question is: “Can a particular PPG algorithm distinguish a true change in arterial pressure from the many other physiological changes that alter the peripheral pulse waveform—and can it do so accurately enough for the intended use?” That is the standard by which PPG-based blood-pressure technology should ultimately be judged.
See also - Does Hilo Core Track Blood Pressure Accurately Across the Day?