Should I log-transform individual timepoint measurements or the absolute change score when my outcome is volume (cm³)?
Should I log-transform individual timepoint measurements or the absolute change score when my outcome is volume (cm³)?
Loading saved threads...
AEP · External communityPost link
External question — Cross Validated Stack Exchange
Author: AEP
Original post: https://stats.stackexchange.com/questions/669333
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I have a longitudinal dataset where my outcome is white matter hyperintensity (WMH) volume measured in cubic centimeters (cm³) from brain MRI, collected at baseline and follow-up for each participant.
I want to analyze the change in volume between the two timepoints. I’m aware that log-transformations are often used to normalize skewed data, stabilize variance, or make proportional changes more interpretable. However, I’m not sure whether the log transformation should be applied to:
Each individual measurement
(baseline and follow-up) before computing the change score, i.e.:
Δ_log = log(Follow-up) - log(Baseline)
This would represent the log of the ratio (fold change) between follow-up and baseline.
The absolute change score
, i.e.:
Δ = Follow-up - Baseline
log(Δ)
This would be a log-transformed version of the raw difference in cm³.
Some additional details:
All volume values are positive, but some changes between baseline and follow-up could be negative (i.e., volume decrease); but perhaps adding a constant to all values could overcome this.
My interest is mainly in comparing the magnitude of change across participants, but it’s not clear whether proportional change or absolute change is the more appropriate scale.
I’m concerned about interpretability, statistical assumptions (normality, homoscedasticity), and whether rankings of change might differ between the two approaches.
Example dataset:
Participant
Baseline (cm³)
Follow-up (cm³)
Raw change
Log of each first (Δ_log)
Log of change
A
10
15
+5
0.405
log(5) = 1.609
B
50
55
+5
0.095
log(5) = 1.609
As shown above, the two approaches give very different interpretations and rankings.
Question:
In this context, which approach is more appropriate — log-transforming the individual measurements before calculating change, or log-transforming the absolute change score — and why? What are the statistical and interpretational trade-offs?
Quote
Report
EdM · External communityPost link
External answer — Cross Validated Stack Exchange
Author: EdM
Original post: https://stats.stackexchange.com/a/669438
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Summarizing comments into an answer:
log-transformations are often used to normalize skewed data, stabilize variance, or make proportional changes more interpretable.
That doesn't necessarily mean that you should log-transform your data, as Peter Flom indicates. If you only have 2 time points and only 1 measurement per participant at each time point, it might be preferable to work in the original scale and use the differences in that scale as the outcome. For example, if you were to perform paired
t
-tests for differences between two groups with your data, it's the distributions of those
paired differences
around the group means that matter with respect to the assumptions underlying the test. The distributions of raw observations don't matter.
If you have measurements on multiple brain regions per participant at each time point, you might prefer to work in the log scale to highlight fractional instead of absolute changes. That decision should be based on your understanding of the subject matter.
which approach is more appropriate — log-transforming the individual measurements before calculating change, or log-transforming the absolute change score — and why?
If you do choose to work with log-transformed data, you certainly should log-transform the original values instead of their differences. As many comments say, there's no reliable way to work with log transforms on potentially negative values. If you log-transform the original values, the differences are then the logs of their ratios. Those have simple interpretations, unlike any of the work-arounds you suggest for transforming the differences.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: AEP Source score (net votes, not local likes): 1 Original post: https://stats.stackexchange.com/questions/669333 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I have a longitudinal dataset where my outcome is white matter hyperintensity (WMH) volume measured in cubic centimeters (cm³) from brain MRI, collected at baseline and follow-up for each participant. I want to analyze the change in volume between the two timepoints. I’m aware that log-transformations are often used to normalize skewed data, stabilize variance, or make proportional changes more interpretable. However, I’m not sure whether the log transformation should be applied to: Each individual measurement (baseline and follow-up) before computing the change score, i.e.: Δ_log = log(Follow-up) - log(Baseline) This would represent the log of the ratio (fold change) between follow-up and baseline. The absolute change score , i.e.: Δ = Follow-up - Baseline log(Δ) This would be a log-transformed version of the raw difference in cm³. Some additional details: All volume values are positive, but some changes between baseline and follow-up could be negative (i.e., volume decrease); but perhaps adding a constant to all values could overcome this. My interest is mainly in comparing the magnitude of change across participants, but it’s not clear whether proportional change or absolute change is the more appropriate scale. I’m concerned about interpretability, statistical assumptions (normality, homoscedasticity), and whether rankings of change might differ between the two approaches. Example dataset: Participant Baseline (cm³) Follow-up (cm³) Raw change Log of each first (Δ_log) Log of change A 10 15 +5 0.405 log(5) = 1.609 B 50 55 +5 0.095 log(5) = 1.609 As shown above, the two approaches give very different interpretations and rankings. Question: In this context, which approach is more appropriate — log-transforming the individual measurements before calculating change, or log-transforming the absolute change score — and why? What are the statistical and interpretational trade-offs?
Checking account access…