Showing posts with label odds. Show all posts
Showing posts with label odds. Show all posts

Friday, March 16, 2018

Likelihood Ratios for Amelia Earhart

Early in the morning of July 2, 1937, Amelia Earhart took off in her twin-engine Lockheed 10E Electra from Papua New Guinea. She was in the midst of a sensational attempt to circle the globe. Her immediate destination was Howland Island. She planned to refuel at this tiny mound. She never made it. The wreckage of her plane and the remains of either Earhart or the navigator who accompanied her have never been located. Or have they?

Earhart probably ran out of fuel while searching for Howland Island. But where? Nikumaroro Island lies some 350 nautical miles south of Howland. Bones found on this uninhabited coral ring in 1939 and sent to Fiji for examination (and now lost) were measured in 1941 by a physician who concluded that they probably belonged to a 45- to 55-year-old male who was around five-and-a-half feet tall.

However, an early release of an article by University of Tennessee anthropologist Richard Jantz, in the University of Florida's new journal, Forensic Anthropology, rejects this conclusion. Arguing that that the skeletal features for sex determination at the time were only weakly probative, Professor Jantz maintains that in 1941, the examiner “could easily have been presented with morphology that he considered male, even though it may have been female.” 1/

As to the lengths of the bones, Dr. Jantz compares estimates for Earhart (inferred from photographs and clothing) with the reported lengths of the Nikumaroro bones using a statistic known as the Mahalanobis distance. By this distance measure for multivariate data, Earhart is more similar to the Nikumaroro bones than 99% of individuals in a sample of bones from 2,776 other people. The sample is not clearly described in the article, but it is well known to forensic anthropologists, and “Jantz told Fox News that 2,776 individuals used in the reference group were all Americans of European ancestry [who] lived during the last half of the 19th century and most of the 20th century ... .” 2/

With respect to this sample, Jantz wrote that:
Earhart’s rank is 19, meaning that 2,758 (99.28%) individuals have a greater distance from the Nikumaroro bones than Earhart, but only 18 (0.65%) have a smaller distance. The rank is subject to sampling variation, so I conducted 1,000 bootstraps of the 2,776 distances, omitting Earhart, then replacing her to determine her rank. Her rank ranged from 9 to 34, the 95% confidence intervals ranging from 12 to 29. If we take the maximum rank resulting from 1,000 bootstraps, 98.77% of the distances are greater and only 1.19% are smaller. If these numbers are converted to likelihood ratios as described by Gardner and Greiner (2006), one obtains 154 using her rank as 19, or 84 using the maximum bootstrap rank of 34. The likelihood ratios mean that the Nikumaroro bones are at least 84 times more likely to belong to Amelia Earhart than to a random individual who ended up on the island.
Let's not get bogged down in the details of the distance measure, bootstrapping, and the conversion to a likelihood ratio. 3/ There is a broader point to note. A likelihood ratio of at least 84 hardly means that “the Nikumaroro bones are at least 84 times more likely to belong to Amelia Earhart than to a random individual” — or even to a random Caucasian-American of the relevant time period. Likelihoods are measures of statistical support for a claim like the one that the particular bones are those of Amelia Earhart. The ratio indicates how much the bone-length data support the Earhart source hypothesis as contrasted to the support those data provide for random-Caucasian-American source hypothesis.

Relative support contributes to the odds in favor of a hypothesis, but it does not express those odds directly. Suppose I pick a coin at random from a box that contains one trick coin (it has heads on both sides) and 128 fair coins. I flip the coin seven times and observe seven heads. The likelihood ratio for the hypotheses of the trick coin as opposed to a fair coin is the probability of the data (seven out of seven heads) for the trick coin divided by the probability for a fair coin. The value of ratio therefore is 1 / (1/2)7 = 1 / (1/128) = 128. But the odds that it is a trick coin are nowhere near 128 to 1. They are 1 to 1. Because it is no more probable that the coin is a trick coin than a fair one, I cannot even say that the preponderance of the statistical evidence favors the trick-coin hypothesis. See Box 1.
Box 1. The Odds on the Trick Coin

Consider what would happen if we repeated the coin picking and tossing experiment 129 times (replacing the picked coin each time). We expect to pick the trick coin from the box only once and hence to see seven heads for that reason one time in 129. We expect to pick fair coins the other 128 times. We also expect that one of the 128 seven-flip tests of these fair coins will produce seven heads. Observing seven heads does not prove that the coin is more likely to be a trick coin than a fair one. We should post the same odds on each hypothesis about the coin. (This heuristic argument easily could be replaced with a more rigorous proof using Bayes' rule.)


Unfortunately, the statement that “the Nikumaroro bones are at least 84 times more likely to belong to Amelia Earhart than to a random individual” sounds like an assertion that the odds on Earhart as opposed to “a random individual” are at least 84 to 1. Mathematically, if there were no other possibilities to consider, these posterior odds would mean that the probability that the bones are Earhart’s is at least 84/(84+1) = 0.988, or just about 99%. If the probability that the bones are from a member of some other ancestral group (such as Micronesians) is, say, 1/10, then the probability for Earhart would decline to 0.889 (about 89%). See Box 2.
Box 2. Going from the Odds of Two Non-exhaustive Hypotheses to a Probability

Let E denote the event that the bones are Earhart’s, let R be the event that they are “a random individual” (among all Caucasian Americans of the time) and let X be the event that they come from someone else. Also, let p, r, and x be the probabilities of each of these events, respectively. Then 84 = p/r, and p+r+x = 1. Substituting and rearranging terms, tt follows that p = 84(1–x) / 85. If x = 0, the probability of E is p = 84/85 = 0.988. If x = 1/10, p = (9/10)×(84/85) = 0.889. 4/
Journalists have picked up on the 99% figure — in some strange ways. The BBC announced that the Forensic Anthropology article “claims they have a 99% match, contradicting an earlier conclusion.” 5/ But being in the upper percentile on a list of distances is not the same as being 99% similar. Fox News crowed  that “Amelia Earhart Disappearance '99 Percent' Solved,” 6/ whatever that may mean.

Other than seeming to conflate the likelihood ratio of 84 with the odds in favor of Earhart, the article does not contend that the probability that the bones are Earhart's is at least 99%. Hewing to the proper interpretation of likelihood as a measure of support for a hypothesis, Professor Jantz wrote that his “analysis ... strongly supports the conclusion that the Nikumaroro bones belonged to Amelia Earhart.” He did go on to mention Bayes' rule and to illustrate what the posterior probability of this hypothesis might be, but following the recommended forensic practice, he did not settle on a specific prior probability. 7/ That probability turns on the non-anthropological information in the case — things like the fact that the Coast Guard cutter off of Howard Island received radio transmissions from Earhart (suggesting that she was not near Nikumaroro Island). Nevertheless, the article seems to propose that the bone lengths in and of themselves prove that the remains probably are Earhart's. It states that:
If [the] sex estimate, can be set aside, it becomes possible to focus attention on the central question of whether the Nikumaroro bones may have been the remains of Amelia Earhart. There is no credible evidence that would support excluding them. On the contrary, there are good reasons for including them. The bones are consistent with Earhart in all respects we know or can reasonably infer. Her height is entirely consistent with the bones. The skull measurements are at least suggestive of female. But most convincing is the similarity of the bone lengths to the reconstructed lengths of Earhart’s bones. Likelihood ratios of 84–154 would not qualify as a positive identification by the criteria of modern forensic practice, where likelihood ratios are often millions or more. They do qualify as what is often called the preponderance of the evidence, that is, it is more likely than not the Nikumaroro bones were (or are, if they still exist) those of Amelia Earhart. If the bones do not belong to Amelia Earhart, then they are from someone very similar to her. And, as we have seen, a random individual has a very low probability of possessing that degree of similarity.
Contrary to the highlighted text, likelihood ratios of 84–154 do not necessarily mean that evidence satisfies the preponderance-of-the-evidence or more-probable-than-not standard of most civil litigation. The legal standard applies to a posterior probability, not to a likelihood ratio standing alone. 8/ Even a large likelihood ratio may not suffice to overcome a small prior probability. (That is what happened in the coin-flipping example of Box 1.) Conversely, even a small likelihood ratio may be enough to boost a prior probability into the more-probable-than-not range.

Professor Jantz recognizes that no one can resolve the historical mystery on the basis of his statistical analysis alone. He writes:
Ideally in forensic practice a posterior probability that remains belong to a victim can be obtained. Likelihood ratios can be converted to posterior odds by multiplying by the prior odds. For example, if we think the prior odds of Amelia Earhart having been on Nikumaroro Island are 10:1, then the likelihood ratios given above become 840–1,540, and the posterior probability is 0.999 in both cases. The prior odds or prior probability pertain to information available before skeletal evidence is considered. It is often impossible to assign specific numbers to the prior probability, because it depends on how the non-osteological evidence is evaluated, and different people will usually evaluate it differently. In jury trials, experts are often advised to testify only to the likelihood ratio developed from the biological evidence. The jury then supplies its own prior odds based on the entire context (e.g., Steadman et al. 2006).
Judging the entire historical record, Jantz adopts a high prior probability (perhaps higher than the 10:1 figure for the prior odds quoted above) to conclude that “[u]ntil definitive evidence is presented that the remains are not those of Amelia Earhart, the most convincing argument is that they are hers.” In other words, the product of the moderately large likelihood ratio and the prior odds (already sufficient to establish a preponderance) is so large that only definitive evidence for an alternative hypothesis could possibly overcome it.

* * *

In the end, do “modern methods produce results that suggest a 99 percent certainty that the bones belonged to Earhart,” as a respected fact-checking website concluded? 8/ Well, the “modern methods” try to exploit the anthropological data more fully than the earlier analyses, but any conclusion about “the certainty that the bones belonged to Earhart” necessarily rests on a judgment of other information as well — the radio transmissions received by the Coast Guard cutter, the failure to spot any signs of Earhart’s presence in a contemporaneous search of the island, other artifacts found on the island in later investigations, and much more.

Professor Jantz is widely reported to have a personal probability of 99% for Earhart as the source of the remains. ABC News, for example, quoted him as stating that "I am 99 percent sure that these bones belong to Amelia Earhart." 10/ This level of belief may be appropriate, based on his review of all the historical information and his latest statistical analysis of the bone lengths. But the article, at least, does not assign a posterior probability to what it presents as “the most convincing argument,” and a forensic anthropologist who did so in court would be relying on knowledge from outside the realm of forensic osteology.

NOTES
  1. Richard L. Jantz, Amelia Earhart and the Nikumaroro Bones: A 1941 Analysis versus Modern Quantitative Techniques, 1(2) Forensic Anthropology 1-16 (2018), available at http://journals.upress.ufl.edu/fa/article/view/525/518.
  2. James Rogers, Amelia Earhart Mystery Solved? Scientist '99 Percent' Sure Bones Found Belong to Aviator, Fox News, Mar. 7, 2018, http://www.foxnews.com/science/2018/03/07/amelia-earhart-mystery-solved-scientist-99-percent-sure-bones-found-belong-to-aviator.html.
  3. The length estimates for Earhart's bones are not exact, bootstrapping is not the same as drawing repeated probability samples from the desired population, and (as discussed in the article) deriving a likelihood ratio involves categorizing continuous measurements into discrete intervals whose size is somewhat arbitrary. Accounting for these sources of uncertainty would produce a broader range of plausible likelihood ratios.
  4. Jantz argues that the sizes of the recovered bones are less typical of Pacific Islanders than of “Euro-Americans,” but the article does not maintain that the probability of a different ancestry is zero or that it should be ignored.
  5. Amelia Earhart: Island Bones 'Likely' Belonged to Famed Pilot, BBC News, Mar. 8, 2018, http://www.bbc.com/news/world-us-canada-43323944.
  6. Rogers, supra note 2.
  7. Cf. Ira M. Ellman & David H. Kaye, Probabilities and Proof: Can HLA and Blood Test Evidence Prove Paternity?, 55 NYU L. Rev. 1131 (1979), available at http://ssrn.com/abstract=1466964.
  8. See, e.g., John Kaplan, Decision Theory and the Factfinding Process, 20 Stan. L. Rev. 1065 (1968); David H. Kaye, Clarifying the Burden of Persuasion: What Bayesian Decision Rules Do and Do Not Do, 3 Int'l J. Evid. & Proof 1 (1999), available at http:/ssrn.com/abstract=2702990
  9. Alex Kasprak, Have Amelia Earhart’s Remains Been Located?, Mar. 15, 2018, https://www.snopes.com/fact-check/amelia-earharts-remains-located/.
  10. E.g., Professor Believes Bones Found on Pacific Island Belong to Amelia Earhart, ABC7 Eyewitness News, http://abc7chicago.com/science/professor-believes-bones-found-on-pacific-island-belong-to-amelia-earhart/3190174/.

Wednesday, May 31, 2017

A Few Statistical and Legal Ideas About the Weight of Evidence

The expression “weight of evidence” has become popular among theorists of forensic science, where it is used to indicate the extent to which findings support the claim that two similar traces originated from the same source as opposed to the claim that they originated from different sources. Speaking more broadly, the idea is that the degree of corroboration a body of evidence provides for a theory or hypothesis depends on the probability of the evidence given that hypothesis compared to the probability of the evidence given other hypotheses. This notion has a rich intellectual history in philosophy, law, and statistics.A recent book review* discusses ways to quantify this measure of corroboration and the motivations for them. Some excerpts follow:

The “likelihood ratio” is a concept that pervades statistics. 31/ As [Richard] Lempert argued, it can be used to define whether an item of evidence is [logically] relevant. For example, in the 1990s researchers developed a prostate cancer test based on the level of prostate-specific antigen (“PSA”). The test, they said, was far from definitive but still had diagnostic value. Should anyone have believed them? A straightforward method for validation is to run the test on subjects known to have the disease and on other subjects known to be disease-free. The PSA test was shown to give a positive result (to indicate that the cancer was present) about 70% of the time when the cancer was, in fact, present, and about 10% of the time when the cancer was not actually present. Thus, the test has diagnostic value. The doctor and patient can understand that positive results arise more often among patients with the disease than among those without it.

But why should we say that the greater probability of the evidence (a positive test result) among cancer patients than among cancer-free patients makes the test diagnostic of prostate cancer? There are three answers. One is that if we use it to sort patients into the two categories, we will (in the long run) do a better job than if we use some totally bogus procedure (such as flipping a coin). This is a “frequentist” interpretation of diagnostic value.

A second justification takes the notion of “support” for a hypothesis as fundamental. 37/ Results that are more probable under a hypothesis H1 about the true state of affairs are stronger evidence for H1 than for any alternative (H2) under which they are less probable. If the evidence were to occur with equal probability under both states, however, the evidence would lend equal support to both possibilities. In this example, such evidence would provide no basis for distinguishing between cancer-free and cancer-afflicted patients. It would have no diagnostic value, 38/ and the test should be kept off the market. The coin-flipping test is like this. A head is no more or less probable when the cancer is present than when it is absent.

A difference between the “frequentist,” long-run justification and the “likelihoodist,” support-based understanding is that the latter applies even when we do not perform or imagine a long series of tests. If it really is more probable to observe the data under one state of affairs than another, it would seem perverse to conclude that the data somehow support the latter over the former. The data are “more consistent” with the state of affairs that makes their appearance on a single occasion more probable (even without the possibility of replication).

The same thing is true of circumstantial evidence in law. Circumstantial evidence E that is just as probable when one party’s account is true as it is when that account is false has no value as proof that the account is true or false. It supports both states of nature equally and is logically irrelevant. To condense these observations into a formula, we can write:
E is irrelevant (to choosing between H1 and H2) if P(E|H1) = P(E|H2),
where P(E|H1) and P(E|H2) are the probabilities of the evidence conditional on (“given the truth of,” or just “given”) the hypotheses. The conditional probabilities (or quantities that are directly proportional to them) have a special name: likelihoods. So a mathematically equivalent statement is that
E is irrelevant if the likelihood ratio L = P(E|H1) / P(E|H2) = 1.
A fancier way to express it is that E is irrelevant if the logarithm of L is 0. Such evidence E has zero “weight” when placed on a metaphorical balance scale that aggregates the weight of the evidence in favor of one hypothesis or the other. 39/ In this [prostate cancer] case, the likelihood ratio for a positive test result is 70% ÷ 10% = 7, which is greater than 1. Thus, the test is relevant evidence in deciding whether the patient has cancer. ...

Nothing that I have said so far involves Bayes’s rule. “Likelihood” and “support” are the primitive concepts. Lempert argued for a likelihood ratio of 1 as the defining characteristic of relevance by relying on a third justification—the Bayesian model of learning. How does this work? Think of probability as a pile of poker chips. Being 100% certain that a particular hypothesis about the world is correct means that all of the chips sit on top of that hypothesis. Twenty-five percent certainty means that 25% of the chips sit on the same hypothesis, and the remaining 75% are allocated to the other hypotheses. 42/ To keep things as simple as possible, let’s assume there are only two hypotheses that could be true. To be concrete, let’s say that H1 asserts that the individual has cancer and that H2 asserts that he does not. Assume that doctors know that men with this patient’s symptoms have a 25% probability of having prostate cancer. We start with 25% of the chips on hypothesis 1 (H1: cancer) and 75% on the alternative (H2: some other cause of the symptoms). Learning that the PSA test is positive for cancer requires us to take some of the chips from H2 and put them on H1. Bayes’s rule dictates just how many chips we transfer. The exact amount generally depends on two things: the percentage of chips that were on H1 (the prior probability) and the likelihood ratio L in this simple situation. ... [T]he very simple structure of Bayes’s rule in this case [is]
Odds(H1) · L = Odds(H1|E).
The rule requires updating the “prior odds” (on the left-hand side) by multiplying by the Bayes factor (which also is the likelihood ratio L) to arrive at the “posterior odds” (on the right-hand side). ...

The crucial point is that multiplication by L = 1 never changes the prior odds. Evidence that is equally probable under each hypothesis produces no change in the allocation of the chips—no matter what the initial distribution. Prior odds of 1:3 become posterior odds of 1:3. Prior odds of 10,000:1 become posterior odds of 10,000:1. The evidence is never worth considering. Again, we can get fancy and place the odds and the likelihood ratio on a logarithmic scale. Then the posterior log odds are the prior log odds plus the weight of the evidence (WOE = log-L):
New LO = Prior LO + WOE. 44/
Evidence that has zero weight (L = 1, log-L = 0) leaves us where we started. Evidence E that does not change the odds (and, hence, the corresponding probability) is uninformative—it is irrelevant. Inversely, evidence that does change the probability is relevant—as [Federal] Rule [of Evidence] 401 states in near-identical terms. This, in a nutshell, is the Bayesian explanation of the rule as it applies to circumstantial evidence. It tracks the text of the rule better than the likelihoodist, support-based analysis, but both lead to the conclusion that relevance vel non turns on whether the likelihood ratio departs from 1. ...

... The simple likelihood ratio is the basic measure that dominates the forensic science literature on evaluative conclusions. However, most writers in this area construe the likelihood ratio as the ratio of posterior odds to prior odds and base its use on that purely Bayesian interpretation. Greater clarity would come from using the related term “Bayes factor” when this is the motivation for the ratio. 51/ [Note 51: The choice of words is not merely a labeling issue. In simple situations, the Bayes factor and the likelihood ratio are numerically equivalent, but more generally, there are conceptual and operational differences. For instance, simple likelihood ratios can be used to gauge relative support within any pair of hypotheses, even when the pair is not exhaustive. But when there are many hypotheses, the Bayes factor is not so simple. See [Peter M. Lee, Bayesian Statistics 140 (4th ed. 2012)], at 141–42. It becomes the usual numerator divided by a weighted sum of the likelihoods for each hypothesis. The weights are the probabilities (conditional on the falsity of the hypothesis in the numerator). For an example, see Tacha Hicks et al., A Framework for Interpreting Evidence, in Forensic DNA Evidence Interpretation 37, 63 (John S. Buckleton et al. eds., 2d ed. 2016). Furthermore, there is disagreement over the use of a likelihood ratio for highly multidimensional data (such as fingerprint patterns and bullet striations) and whether and how to express uncertainty with respect to the likelihood ratio itself. Compare Franco Taroni et al., Dismissal of the Illusion of Uncertainty in the Assessment of a Likelihood Ratio, 15 Law, Probability & Risk 1, 2 (2016), with M.J. Sjerps et al., Uncertainty and LR: To Integrate or Not to Integrate, That’s the Question, 15 Law, Probability & Risk 23, 23–26 (2016). ... ] ...

The obvious Bayesian measure of probative value is the Bayes factor (B). In the examples used here, B is equal to the likelihood ratio L, and therefore the statisticians’ “weight of evidence” is WOE = log-B = log-L. 58/ The value of L in these cases tells us just how much more the evidence supports one theory than another and hence—this is the Bayesian part—just how much we should adjust our belief (expressed as odds) for any starting point. For the PSA test for cancer, L = 7 is “the change in odds favoring disease.” A test with greater diagnostic value would have a larger likelihood ratio and induce a stronger shift toward that conclusion. ... [T]he likelihood-ratio measure (or variations on it), which keeps prior probabilities out of the picture, is more typically used to describe the value of test results as evidence of disease or other conditions in medicine and psychology. Using the same measure in law has significant advantages. ...

NOTES
* David H. Kaye, Digging into the Foundations of Evidence Law, 115 Mich. L. Rev. 915 (2017)  (reviewing The Michael J. Saks & Barbara A. Spellman, Psychological Foundations of Evidence Law (2016)).

31. Vic Barnett, Comparative Statistical Inference 306 (3d ed. 1999) (“The principles of maximum likelihood and of likelihood ratio tests occupy a central place in statistical methodology.”); see, e.g., id. at 178–80 (describing likelihood ratio tests in frequentist hypothesis testing); N. Reid, Likelihood, in Statistics in the 21st Century 419 (Adrian E. Raftery et al. eds., 2002).

37. A “support function” can be required to have several appealing, formal properties, such as transitivity and additivity. E.g., A.W.F. Edwards, Likelihood 28–32 (Johns Hopkins Univ. Press, expanded ed. 1992) (1972). It also can be derived, in simple cases, from other, arguably more fundamental, principles. E.g., Barnett, supra note 31, at 310–11.

39. See generally I. J. Good, Weight of Evidence and the Bayesian Likelihood Ratio, in the Use of Statistics in Forensic Science 85 (C.G.G. Aitken & D.A. Stoney eds., 1991); I. J. Good, Weight of Evidence: A Brief Survey, in 2 Bayesian Statistics 249 (J.M. Bernardo et al. eds., 1985) (providing background information regarding the use of Bayesian statistics in evaluating weight of evidence). ...


42. If the individual were to keep some of the chips in reserve, the analogy between the fraction of them on a hypothesis and the kind of probability that pertains to random events such as games of chance would break down.

44. A deeper motivation for using logarithms may lie in information theory, but, if so, it is not important here. See Solomon Kullback, Information Theory and Statistics (1959).

58. The logarithm of B has been called “weight of evidence” since 1878. I.J. Good, A. M. Turing’s Statistical Work in World War II, 66 Biometrika 393, 393 (1979) ... . While working in the town of Banbury to decipher German codes, Alan Turing famously (in cryptanalysis and statistics, at least) coined the term “ban” to designate a power of 10 for this metaphorical weight. Good, supra, at 394. Thus, a B of 10 is 1 ban, 100 is 2 ban, and so on.

Sunday, August 31, 2014

Hazard Ratios and Heart Failure

Today’s big news in medicine is a new drug, designated LCZ696 by its manufacturer, Novartis. According to the New York Times, LCZ696 “has shown a striking efficacy in prolonging the lives of people with heart failure and could replace what has been the bedrock treatment for more than 20 years.” [1] Specifically, more than 8,400 patients in 47 countries enrolled in a randomized, double-blind experiment in which they received either LCZ696 or an ACE inhibitor called enalapril (in addition to whatever else their doctors prescribed).

The trial was halted after a median follow-up time of 27 months “ because the boundary for an overwhelming benefit with LCZ696 had been crossed.” [2] “By that point, 21.8 percent of those who received LCZ696 had died from a cardiovascular cause or had been hospitalized for worsening heart failure. That figure was 26.5 percent for those receiving enalapril. That represents a 20 percent relative reduction in risk using a statistical measure called the hazard ratio.” [1]

This is good news for patients (if the drug receives regulatory approval and performs as expected in practice). But the account in the Times poses a small statistical puzzle. How does the difference between 21.8 and 26.5 percentage points translate into “a 20 percent relative reduction in risk”? The average risk across patients dropped by 26.5 – 21.8 = 4.7 percentage points. This absolute reduction is appreciable, but 4.7 percentage points is not 20% of the original 26.5 percent risk of hospitalizations or deaths in the control group (4.7 / 26.5 = 17.7%). What accounts for the discrepancy?

The answer lies in the details of a technique known in biostatistics as survival analysis. The statistical technique is not limited to the analysis of death rates. It can be applied to all sorts of situations involving different times to some outcome. The outcome can be the overruling of a Supreme Court case, the firing of a worker, or the exoneration of a prison inmate sentenced to die, to pick a few examples from forensic statistics.

So what does the 20% “relative reduction in risk” cited in the Times article mean? Well, a hazard function is the probability that if you survive to a given time t (the event in question has not already occurred), you will survive in the next instant. A hazard ratio is the ratio of the hazard in the treatment group to the hazard in the control group at t. The heart failure study used an estimation procedure known as proportional hazards regression, which assumes that the hazard in one group is a constant proportion of the hazard in the other group. Under this assumption, in a clinical trial where death is the endpoint, the hazard ratio indicates the relative likelihood of death in treated versus control subjects at any given point in time.

Thus, unlike the ordinary relative risk discussed in many court opinions, the “hazard ratio” is not simply the proportion with a disease in an exposed group divided by the proportion in an unexposed group. In the LCZ696 study, the hazard ratio was 0.80, meaning that the probability that a randomly selected patient taking LCZ696 would die from or be hospitalized for heart failure the next day is 80% of the probability for a randomly selected patient taking enalapril. To put it another way, the probability of hospitalization or death tomorrow from heart failure drops by 20% when LCZ696 is substituted for enalapril.

Yet a third formulation is that the odds that a randomly selected patient treated with LCZ696 will be hospitalized or die sooner than a randomly selected control patient are 0.8 (to 1) — that's 4 to 5, corresponding to a probability of 4/9 = 44%. [3]

How long either patient can expect to live and avoid hospitalization from heart failure is another story. As one article on hazard ratios explains, “[t]he difference between hazard-based and time-based measures is analogous to the odds of winning a race and the margin of victory.” [3] By itself, the hazard ratio picks the winning horse (probably), but it does not give the number of lengths for its expected success.

References
  1. Andrew Pollack, New Novartis Drug Effective in Treating Heart Failure, N.Y. Times, Aug.31, 2014, at A4
  2. John J.V. McMurray et al., Angiotensin–Neprilysin Inhibition versus Enalapril in Heart Failure, New Engl. J. Med., Aug. 30, 2014
  3. Spotswood L. Spruance et al., Hazard Ratio in Clinical Trials, 48 Antimicrobial Agents and Chemotherapy 2787 (2004)