Thursday, April 2, 2020

"He (or She) Tested Negative (or Positive) for Coronavirus"

The news is full of stories of celebrities and public officials who tested positive or negative "for coronavirus" or "for COVID-19."
  • Prime Minister Boris Johnson has tested positive for coronavirus -- BBC News 3/27/20
  • Rapper Scarface revealed he has tested positive for coronavirus -- The Daily Beast 3/26/20
  • Rep. Mike Kelly (R-Pa.) announced Friday he has tested positive for the coronavirus -- The Hill, 3/27/20
  • An Arizona State University professor said he tested positive for COVID-19 -- KTAR 3/27/20
  • Trump tested negative for coronavirus -- CNN, 3/14/20
  • Charles Barkley announced he tested negative for the coronavirus -- USA Today, 3/23/20
  • Romney says he tested negative for coronavirus -- The Hill, 3/24/20
  • Lindsey Graham says he tested negative for coronavirus -- CNN, 3/15/20
  • Ayanna Pressley tests negative for COVID-19 -- CNN, 3/27/20
What can anyone really conclude from a negative or positive finding? How well do these findings answer the question of whether someone is infectious, or ill because of an infection? This posting seeks to explain why convincing estimates of test sensitivity and specificity are hard to come by. It also sketches the kind of additional reasoning that would be necessary to supply estimates of the probability a person is infected with the virus or ill from COVID-19 in light of the test results. (I am outside my comfort zone in parts of this posting -- corrections are welcome.)

Tests for What?

To begin with, we need to distinguish between the disease -- Coronavirus Disease 2019 (COVID-19) -- and the virus itself -- Sudden Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2). The virus hijacks the machinery of human cells to replicate itself. Initially, it tends to reside in the mucous membranes of the upper nose and throat, but in more serious cases, it moves from the upper respiratory tract to the lungs. The disease spreads primarily through respiratory droplets from an infected individual that end up in the mouth, nose, or eyes of another person.

In a sense, the tests in the news are not tests for the disease -- even though many of their creators call them tests for COVID-19. \1/They are molecular diagnostics tests for the presence of certain sequences (of the nucleotide base-pairs) that are characteristic of SARS-CoV-2. \2/ If these sequences are detected in a swab from the person under investigation (a PUI), the test is said to be positive. If these sequences are not detected, the result is negative.

Operating Characteristics: Sensitivity and Specificity

In an ideal test for being infected with the virus, positives would only arise when SARS-CoV-2 is present in the PUI, and negatives would only occur when it is not. The probability of a positive result (+) given that the PUI harbors the specific SARS virus then would be 1. As we will see, this probability is not precisely known, but it surely is less than 1. If we let S2 stand for the event that the PUI has the virus SARS-Cov-2, we can write this conditional "true positive" probability, or test sensitivity, as Pr(+ | S2). The "|" in the expression is read "given" or "conditional on."

One other probability is needed to characterize the accuracy of the test. The specificity indicates how accurately the test indicates that a PUI does not harbor the virus. Ideally, the specificity, Pr(− | not-S2), also is  1. That is, whenever the PUI is not infected, the test is negative. But, once again, no real-world test for infection performs this well.

To see why, the following path diagram for test results may be helpful:
Figure 1. What might produce positive and negative test results

The diagram shows that a positive test result (TEST +) could be explained either by viruses from the PUI or by contamination on the swab. A negative test result (TEST −) could be explained either by the absence of any infection in the PUI, by an infection that has not generated enough viruses to signal a positive result, or by problems with the chemistry of the test. If any these paths have a nonzero probability, the sensitivity and specificity are less than 1.

Even this list of explanations assumes that "virus" in the diagram refers strictly to the SARS-CoV-2 strain that the test is designed to detect. If another type of coronavirus, a rhinovirus, parainfluenza virus, adenovirus, etc., has sufficient sequence similarity in the few regions tested to be mistaken for SARS-CoV-2, then the similar viruses on the swab could produce a signal. That would make the test even less specific to the infectious agent for COVID-19. Conversely, if other strains of SARS-CoV-2 exist and have sufficiently different sequences in the regions of the virus's genome that the test covers, the test will miss them, reducing its sensitivity. The FDA calls this aspect of sensitivity "inclusivity." \3/

So what are the sensitivity and specificity of the tests that have been released under emergency use authorizations from the FDA? The laboratories that rushed to develop the tests based on the viral genome performed limited experiments to assess (1) how much virus on the swab would be detectable; (2) whether other types of viruses would be detected instead of the real target; and (3) the probabilities implicit in the pathways in the blue boxes in the diagram.

Reported Laboratory Validation Data

To distribute or perform tests, manufacturers or laboratories must file validity studies with the Food and Drug Administration, which insists that "[a]ll clinical tests should be validated prior to use. In the context of a public health emergency, it is especially important that tests are validated as false results can have broad public health impact beyond that to the individual patient." \4/

For example, the Laboratory Corporation of America's Accelerated Emergency Use Authorization (EUA) Summary for its COVID-19 RT-PCR Test explains that "[t]he COVID-19 RT-PCR test is a real-time reverse transcription polymerase chain reaction (rRT-PCR) test for the qualitative detection of nucleic acid from SARS-CoV-2 in upper and lower respiratory specimens (such as nasopharyngeal or oropharyngeal swabs, sputum, lower respiratory tract aspirates, bronchoalveolar lavage, and nasopharyngeal wash/aspirate or nasal aspirate) collected from individuals suspected of COVID-19 by their healthcare provider."

The summary asserts that "SARS-CoV-2 RNA is generally detectable in respiratory specimens during the acute phase of infection" -- in other words, the test is somewhat sensitive to the disease. But that covers a lot of territory. Without data on the probabilities in the path from PUI infected to viruses in the speciment-collection site to a detectable quantity on the swab (and other possible paths), we are left with vague statements in the summary, such as "[p]ositive results are indicative of the presence of SARS-CoV-2 RNA" on the swab or other sample, and "[n]egative results do not preclude SARS-CoV-2 infection."

Of course, the test only is designed to signal the presence of viruses in the sample (the paths in the blue boxes in Figure 1). As I noted at the outset, it does not purport to be a test for the disease itself. How well does this test accomplish its more limited task? Naturally, detection of the virus depends on the quantity of viruses that are actually present. LabCorp and other test developers follow the simplistic approach of defining a fixed "Limit of Detection (LoD)." What, then, is the probability of detection at and above the limit?

According to the summary, "[t]he LoD study established the lowest concentration of SARS-CoV-2 (genome copies (cp)/μL) that can be detected by the COVID-19 RT-PCR test at least 95% of the time." This limit came from creating mock specimens with known quantities of the virus ("spiking the quantified live SARS-CoV-2 into negative respiratory clinical matrices") and reducing the quantity to the point at which 19 out of 20 specimens tested positive. Using only 20 mock specimens for each concentration, \5/ "[t]he study results showed that the LoD of the COVID-19 RT-PCR test is 6.25 cp/μL (19/20 positive)."

Although 19/20 describes the sample data, one cannot be entirely confident that the test really has a 95% sensitivity at the selected concentration for the LoD. Even if the true sensitivity at the 6.25 concentration were, say, 18 out of 20, we would find exactly 19 out of 20 replicates to be positive (as occurred in the LoD study) more than a quarter of the time. \6/ Likewise, the 0.95 sensitivity criterion for the limit of detection could have led to twice the reported LoD in a study with the same sample size. If the detection probability (sensitivity) were 95% at the next level up (12.5 cp/μL in this study), it could well be that all 20 replicates would be positive. The probability of that datum is 36% (0.9520 = 0.358).

Having chosen 6.25 for the LoD concentration, LabCorp proceeded to a "Clinical Evaluation." More precisely, "[a] contrived clinical study was performed." For brevity, I will just describe the results for NP swabs. (The data on BALs were the same.)

No. samplesConcentrationTest −Test +
500500
101×LoD010
102×LoD010
104×LoD010
108×LoD010

For these outcomes, the summary derives the following statistics:
  • "Positive Percent Agreement 40/40 = 100% (95% CI: 91.24% - 100%)"
  • "Negative Percent Agreement 50/50 = 100% (95% CI: 92.87% -100%)"
The first confidence interval is an estimate for the sensitivity, based the 40 positive test swabs pooled over the four geometrically decreasing concentrations. This interval, and the second one, for the specificity, suggest that the test is good at distinguishing between swabs spiked with between 6.25 and 50 cp/μL of the virus, on the one hand, and swabs with no SARS-CoV-2 at all, on the other.

But it is not clear what this sample of 90 tests is representative of. The efficacy of the test in discriminating between a virus-free swab and a virusy one depends on the how many viruses are on the swab. If we contrast the 10 swabs constructed to have the reported limit of detection (6.25) with the 50 with no viruses, the observed sensitivity in the experimental sample is still 1, but because 10 is a small sample size, the 95% Clopper-Pearson CI extends as low as 0.69. The lower end of the interval for the specificity is still 0.93. To discern some sort of average sensitivity and specificity for swabs from patients, one would need to know the distribution of viral concentrations in the patient population.

LabCorp's validity study contains further data on "Analytical Specificity." This is not the specificity for the classification for virus-present versus virus-absent seen in the "Clinical Evaluation." It concerns the possibility that a different virus could generate (false) positive results -- something that, as previously noted, would make the test even less specific. The summary lists bacteria and viruses that did not produce positive test results (in an unspecified number of tests). This is consistent with the fact that a number of them have "no homology with primers and probes of the COVID-19 RT-PCR test." In other words, their nucleic acid sequences are substantially different from the sequences of SARS-CoV-2 used in the test. As such, the SARS-CoV-2 amplication and detection process should not react to the sequences from at least this set of other organisms.

Reporting Test Results Without Quantitative Information

The advice from testing companies and laboratories does not even try to supply estimates of sensitivity and specificity -- either for the diagnosis of COVID-19 or the presence of SARS-CoV-2 on specimens. For example, the Fact Sheet for Healthcare Providers: Labcorp's COVID-19 RT-PRC Test - LabCorp (Mar. 16, 2020) contains the following questions and answers:
What does it mean if the specimen tests positive for the virus that causes COVID-19?
A positive test result for COVID-19 indicates that RNA from SARS-CoV-2 was detected, and the patient is infected with the virus and presumed to be contagious. Laboratory test results should always be considered in the context of clinical observations and epidemiological data ....
LabCorp's COVID-19 RT-PCR Test has been designed to minimize the likelihood of false positive test results. ...
What does it mean if the specimen tests negative for the virus that causes COVID-19?
A negative test result for this test means that SARS-CoV-2 RNA was not present in the specimen above the limit of detection. However, a negative result does not rule out COVID-19 and should not be used as the sole basis for treatment or patient management decisions. A negative result does not exclude the possibility of COVID-19.
When diagnostic testing is negative, the possibility of a false negative result should be considered in the context of a patient’s recent exposures and the presence of clinical signs and symptoms consistent with COVID-19. ...
This advice is oddly phrased and not terribly helpful. Among other things, \7/ does the statement that the test is designed to "minimize" the false-positive probability mean that the test maximizes the sensitivity to the point that its complement, the false-positive probability, is 0? That the test design makes the FPP higher than some alternative designs that were considered? Of course, the positive test "indicates" that the viral RNA is present, but how strong is the indication? \8/ And, how should the test result -- whether positive or negative -- be evaluated "in the context of clinical observations and epidemiological data"? The Fact Sheet leaves the healthcare providers for whom it is written at sea.

It is all but impossible to answer the last two questions without understanding Bayes' rule -- a formula for updating a previously established probability in the light of new information such as a symptom or a test result. Suffice it to say that the probability of COVID-19 in the patient is a function of (1) the prevalence of the disease among persons who are like the patient in their demographic and geographic characteristics and medical histories; (2) the sensitivity and specificity of the symptoms (things like a fever and a cough) in this population; and (3) the sensitivity and specificity of the test for SARS-CoV-2 in this population.

Today's Bottom Line

On the basis of the kind of information collected here, it is safe to say that a positive test result raises the odds of COVID-19 and a negative result lowers them. But by how much? To make better use of the tests in diagnosing COVID-19, their operating characteristics should be measured by validating the tests against specimens from patients who are known to be suffering from COVID-19.

The Wall Street Journal has an alarming statistic for the false negative rate. Its answer to the question "Are tests accurate?" is
  • Health experts say they now believe nearly one in three patients who are infected are nevertheless getting a negative test result. They caution that only limited data are available, and their estimates are based on their own experience in the absence of hard science.
  • That picture is troubling, many doctors say, as it casts doubt on the reliability of a wave of new tests developed by manufacturers, lab companies and the CDC. Most of these are operating with minimal regulatory oversight and little time to do robust studies amid a desperate call for wider testing. \9/
A false-negative rate of 1/3 is the same as a sensitivity of 2/3s. (Proof: Let C stand for "has COVID-19." Then Pr(–|C) + Pr(–|not-C) = false negative rate + sensitivity = 1/3 + 2/3 = 1. This notation makes it clear that now we are speaking of conditional probabilities for the disease rather than the presence of the virus in the specimen at or above the LoD.)

A discussion in the Internet Book of Critical Care \10/ refers to one or two studies along these lines. It suggests that in practice, the sensitivity and specificity are each below 80%:
There are several major limitations, which make it hard to precisely quantify how RT-PCR performs.
  1. RT-PCR performed on nasal swabs depends on obtaining a sufficiently deep specimen. Poor technique will cause the PCR assay to under-perform.
  2. COVID-19 isn't a binary disease, but rather there is a spectrum of illness. Sicker patients with higher viral burden may be more likely to have a positive assay. Likewise, sampling early in the disease course may reveal a lower sensitivity than sampling later on.
  3. Most current studies lack a “gold standard” for COVID-19 diagnosis. For example, in patients with positive CT scan and negative RT-PCR, it's murky whether these patients truly have COVID-19 (is this a false-positive CT scan, or a false-negative RT-PCR?). ...
Specificity seems to be high (although contamination can cause false-positive results), [but] sensitivity may not be terrific. ... In a case series diagnosed on the basis of clinical criteria and CT scans, the sensitivity of RT-PCR was only ~70% (Kanne 2/28). Sensitivity varies depending on assumptions made about patients with conflicting data (e.g. between 66-80%) (Ai et al.). ... Among patients with suspected COVID-19 and a negative initial PCR, repeat PCR was positive in 15/64 patients (23%). This suggests a PCR sensitivity of <80%. Conversion from negative to positive PCR seemed to take a period of days, with CT scan often showing evidence of disease well before PCR positivity (Ai et al.).

Bottom line?
PCR seems to have a sensitivity somewhere on the order of ~75%. A single negative RT-PCR doesn't exclude COVID-19 (especially if obtained from a nasopharyngeal source or if taken relatively early in the disease course). If the RT-PCR is negative but suspicion for COVID-19 remains, then ongoing isolation and re-sampling several days later should be considered.

An 80% sensitivity and specificity implies that the test changes the odds of the disease by a factor of only 80/20 = 4. For such a test, if the physician's prior odds (those formed before receiving the test result) were, say, 6:1 in favor of COVID-19, a positive test result would change them to 24:1. The posterior probability is thus 24/25 = 96%. A negative test result would shift the odds from 1:6 for not-COVID-19 to 4:6. These latter odds are equivalent to 6:4 on COVID-19 (a probability of disease of 6/10 = 60%). In short, the starting probability of 6/7 = 87% went up to 96% or down to 60%, depending on whether the test came back positive or negative. If the starting odds were reversed -- 1:6 on COVID-19 prior to the test -- the posterior probabilities of the disease would be lower -- 40% with the positive test result, and only 4% with a negative test result.

By way of comparison, one study of the much maligned technique of microscopic hair comparisons for identity used mitochondrial DNA tests as the gold standard for accuracy. It gave rise to a likelihood ratio for a positive association between the crime-scene hair fibers and the suspects' head hairs of a little under 3. \11/ That is not impressive, but if the estimates proposed by the Critical Care doctors are correct about the tests for SARS-CoV-2, the probative value of a hair association is not all that different from the diagnostic value of a positive molecular diagnostics test.

UPDATE (8/28/20)
A clear discussion of test sensitivity, specificity, and positive predictive value can be found in the International Statistical Institute's blog posting by John Bailar, My COVID-19 Test Is Positive … Do I Really Have It?, Statisticians React to the News, Aug. 25, 2020. It focuses on interpreting rapid antigen screening test results in combination with confirmatory PCR tests of the kind discussed here but does not delve deeply into the estimated sensitivity and specificity of any of the tests. It proposes further further reading on this topic in Lauren Kucirka & Justin Lessler, COVID-19 Story Tip: Beware of False Negatives in Diagnostic Testing of COVID-19, Johns Hopkins Medicine Newsroom, May 26, 2020, ("describing work suggesting false negative rates > 20% for RT-PCR tests and that test accuracy changes over time course of disease"), and Rob Stein, Study Raises Questions About False Negatives From Quick Covid-19 Test, NPR Morning Edition, Apr. 21, 2020 (reporting that "[r]esearchers at the Cleveland Clinic tested 239 specimens known to contain the coronavirus using five of the most commonly used coronavirus tests, including the Abbott ID NOW [which] only detected the virus in 85.2% of the samples, meaning it had a false-negative rate of 14.8 percent.").

NOTES

  1. The names (and other information) on the tests that have received Emergency Use Authorization (EUA) from the FDA are listed at https://www.fda.gov/emergency-preparedness-and-response/mcm-legal-regulatory-and-policy-framework/emergency-use-authorization#2019-ncov.
  2. There also are serological tests. These look for antibodies in the blood. For a description of various types of tests, see Cormac Sheridan, Fast, Portable Tests Come Online to Curb Coronavirus Pandemic, Nature Biotechnology, Mar. 23, 2020.
  3. The FDA's Policy for Diagnostic Tests for Coronavirus Disease-2019 during the Public Health Emergency, Mar. 16, 2020, contains the following "recommendations regarding the minimum testing that should be performed to ensure analytical and clinical validity" for "tests that detect SARS-CoV-2 nucleic acids from human specimens" (pp. 9-10):
    (1) Limit of Detection

    FDA recommends that laboratories document the limit of detection (LoD) of their SARS-CoV-2 assay. FDA generally does not have concerns with spiking RNA or inactivated virus into artificial or real clinical matrix (e.g., Bronchoalveolar lavage [BAL] fluid, sputum, etc.) for LoD determination. FDA recommends that laboratories test a dilution series of three replicates per concentration, and then confirm the final concentration with 20 replicates. For this guidance, FDA defines LoD as the lowest concentration at which 19/20 replicates are positive. If multiple clinical matrices are intended for clinical testing, FDA recommends that laboratories submit in their EUA requests the results from the most challenging clinical matrix to FDA. For example, if testing respiratory specimens (e.g., sputum, BAL, nasopharyngeal (NP) swabs, etc.), laboratories should include only results from sputum in their EUA request.

    (2) Clinical Evaluation

    In the absence of known positive samples for testing, FDA recommends that laboratories confirm performance of their assay with a series of contrived clinical specimens by testing a minimum of 30 contrived reactive specimens and 30 non-reactive specimens. Contrived reactive specimens can be created by spiking RNA or inactivated virus into leftover clinical specimens, of which the majority can be leftover upper respiratory specimens such as NP swabs, or lower respiratory tract specimens such as sputum, etc. We recommend that twenty of the contrived clinical specimens be spiked at a concentration of 1x-2x LoD, with the remainder of specimens spanning the assay testing range. For this guidance, FDA defines the acceptance criteria for the performance as 95% agreement at 1x-2x LoD, and 100% agreement at all other concentrations and for negative specimens.

    (3) Inclusivity

    Laboratories should document the results of an in silico analysis indicating the percent identity matches against publicly available SARS-CoV-2 sequences that can be detected by the proposed molecular assay. FDA anticipates that 100% of published SARS-CoV-2 sequences will be detectable with the selected primers and probes.

    (4) Cross-reactivity

    At a minimum, FDA believes an in silico analysis of the assay primer and probes compared to common respiratory flora and other viral pathogens is sufficient for initial clinical use. For this guidance, FDA defines in silico cross-reactivity as greater than 80% homology between one of the primers/probes and any sequence present in the targeted microorganism. In addition, FDA recommends that laboratories follow recognized laboratory procedures in the context of the sample types intended for testing for any additional cross-reactivity testing.
  4. Id. at 4.
  5. Id. at 9 ("FDA recommends that laboratories test a dilution series of three replicates per concentration, and then confirm the final concentration with 20 replicates.").
  6. In twenty independent tests of swabs with a probability of 18/20 = 0.9 of a detection on each test, the binomial probability of exactly 19 detections is 20 × 0.919 × 0.1 = 0.27.
  7. I think the first sentence is intended to say that a positive test result indicates (with some probability) that the specific RNA is present in the specimen, which implies (with some probability) that the patient is infected with SARS-CoV-2. And, I suspect that the first sentence in the second answer should read, a "negative test result indicates that SARS-CoV-2 RNA was not present in the specimen above the limit of detection" or a "negative test result means that SARS-CoV-2 RNA was not detected in the specimen above the limit of detection."
  8. Despite the warning in the first box about clinical and epidemiologic information, LabCorp apparently regards a positive result on its test as "definitive" in and of itself. A Q&A page on the LabCorp's website does not even include a question about the meaning of a positive result. It only asks whether "a negative result from LabCorp’s testing for COVID-19 mean[s] that a patient is definitely not infected." Its answer is
    Not necessarily. LabCorp’s testing for COVID-19 detects the virus directly, within the established limits of detection for which it was validated. A positive result is considered definitive evidence of infection. However, a negative result does not definitively rule out infection. As with any test, the accuracy relies on many factors:
    • The test may not detect virus in an infected patient if the virus is not being actively shed at the time or site of sample collection.
    • The amount of time that an individual was exposed prior to the collection of the specimen can also influence whether the test will detect the virus.
    • Individual response to the virus can differ.
    • Whether the specimen we receive was collected properly, sent promptly, and packaged correctly. Test results are a critical part of any diagnosis, but must be used by the clinician along with other information to form a diagnosis.
    Q&A, LabCorp's Testing for COVID-19, Mar. 25, 2020, https://www.labcorp.com/assets-media/2330. At least the last bulleted item pertains to positive as well as negative results.
  9. WSJ Staff, Who Has Covid-19? What We Know About Tests for the New Coronavirus?, updated Apr. 2, 2020 7:46 pm ET, https://www.wsj.com/articles/who-has-covid-19-what-we-know-about-tests-for-the-new-coronavirus-11585868185
  10. Josh Farkas, COVID-19, Internet Book of Critical Care, Mar. 2, 2020 (updated Mar. 29, 2020), https://emcrit.org/ibcc/COVID19/.
  11. David H. Kaye, Ultracrepidarianism in Forensic Science: The Hair Evidence Debacle, 72 Wash. & Lee L. Rev. Online, 227–254 (2015) (discussing the study), available at ssrn.com/abstract=2647430.

Sunday, March 15, 2020

SARS-CoV-2 and Risk Perception

Are the governmental, social, and individual responses to the Covid-19 pandemic proportionate to the risks? Given the uncertainties about the SARS-CoV-2 virus and the effectiveness of the responses to its transmission, it is hard to give a definitive answer. Lockdowns are like saturation bombing compared to smart missiles, and public health experts are discussing alternative strategies that rely on rapid tests and the government's ability to track people's movements so as to allow more precise isolations. [1]

All sorts of statistics are important in formulating public policy (and for informing individuals). At the end of February, the New England Journal of Medicine published the following statistics from the China Medical Treatment Expert Group for Covid-19 (derived from 1099 patients with laboratory-confirmed cases of Covid-19 from 552 hospitals in 30 provinces, autonomous regions, and municipalities in mainland China through January 29, 2020) [2]:

Age47 years median
45-58 interquartile range
Female41.9%
Admitted to ICU5%
Invasive mechanical ventilation2.3%
Died1.4%
Had direct contact with wildlife1.9%
Incubation period4 days median
2-7 interquartile range

Generalizing to the fatality rate for the population as a whole is obviously tricky. As the Chinese Expert Group commented,
We found a lower case fatality rate (1.4%) than the rate that was recently reportedly, probably because of the difference in sample sizes and case inclusion criteria. Our findings were more similar to the national official statistics, which showed a rate of death of 3.2% among 51,857 cases of Covid-19 as of February 16, 2020. Since patients who were mildly ill and who did not seek medical attention were not included in our study, the case fatality rate in a real-world scenario might be even lower. Early isolation, early diagnosis, and early management might have collectively contributed to the reduction in mortality in Guangdong.
To supply perspective, the journal offered an essay [3] that observed
History suggests that we are actually at much greater risk of exaggerated fears and misplaced priorities. There are many historical examples of panic about epidemics that never materialized (e.g., H1N1 influenza in 1976, 2006, and 2009). There are countless other examples of societies worrying about a small threat (e.g., the risk of Ebola spreading in the United States in 2014) while ignoring much larger ones hidden in plain sight. SARS-CoV-2 had killed roughly 5000 people by March 12. That is a fraction of influenza’s annual toll. While the Covid-19 epidemic has unfolded, China has probably lost 5000 people each day to ischemic heart disease. So why do so many Americans refuse influenza vaccines? Why did China shut down its economy to contain Covid-19 while doing little to curb cigarette use? Societies and their citizens misunderstand the relative importance of the health risks they face. The future course of Covid-19 remains unclear (and I may rue these words by year’s end). Nonetheless, citizens and their leaders need to think carefully, weigh risks in context, and pursue policies commensurate with the magnitude of the threat. [2]
Of course, these statistics are a small part of all the available data, and far more alarming conclusions about exponential growth (it's present), containment (it's too late), and mitigation ("flatten the curve" to give countries more time to prepare to handle the fraction of cases requiring hospitalization -- around 5%?) can be drawn from modeling the true rate [4]. But even the very high fatality rate in Italy is not easy to interpret. [5]

The lack of reliable data makes it difficult to formulate a credible policy. A provocative essay by John Ioannidis [6] argued that
The data collected so far on how many people are infected and how the epidemic is evolving are utterly unreliable. Given the limited testing to date, some deaths and probably the vast majority of infections due to SARS-CoV-2 are being missed. We don’t know if we are failing to capture infections by a factor of three or 300. Three months after the outbreak emerged, most countries, including the U.S., lack the ability to test a large number of people and no countries have reliable data on the prevalence of the virus in a representative random sample of the general population.

This evidence fiasco creates tremendous uncertainty about the risk of dying from Covid-19. Reported case fatality rates, like the official 3.4% rate from the World Health Organization, cause horror — and are meaningless. Patients who have been tested for SARS-CoV-2 are disproportionately those with severe symptoms and bad outcomes. As most health systems have limited testing capacity, selection bias may even worsen in the near future.

The one situation where an entire, closed population was tested was the Diamond Princess cruise ship and its quarantine passengers. The case fatality rate there was 1.0%, but this was a largely elderly population, in which the death rate from Covid-19 is much higher.

Projecting the Diamond Princess mortality rate onto the age structure of the U.S. population, the death rate among people infected with Covid-19 would be 0.125%. But since this estimate is based on extremely thin data — there were just seven deaths among the 700 infected passengers and crew — the real death rate could stretch from five times lower (0.025%) to five times higher (0.625%). It is also possible that some of the passengers who were infected might die later, and that tourists may have different frequencies of chronic diseases — a risk factor for worse outcomes with SARS-CoV-2 infection — than the general population. Adding these extra sources of uncertainty, reasonable estimates for the case fatality ratio in the general U.S. population vary from 0.05% to 1%.

That huge range markedly affects how severe the pandemic is and what should be done. A population-wide case fatality rate of 0.05% is lower than seasonal influenza. If that is the true rate, locking down the world with potentially tremendous social and financial consequences may be totally irrational. It’s like an elephant being attacked by a house cat. Frustrated and trying to avoid the cat, the elephant accidentally jumps off a cliff and dies.
Would it make sense to isolate the older part of population and allow others to return to school and work? Adjusting "universal quarantines" along these lines has its advocates. [1,7]

[The above was last updated on March 28, 2020.]
Note of March 26, 2020: The single best document for perspective on the virus, the disease,  the response to it, and the available precaution for individuals that I have read is an informal report on "How to fight the coronavirus SARS-CoV-2 and its disease, COVID-19" from Michael Lin.


REFERENCES
  1. Sharon Begley, When Can We Let Up? Health Experts Craft Strategies to Safely Relax Coronavirus Lockdowns, STAT. Mar. 25, 2012, https://www.statnews.com/2020/03/25/coronavirus-experts-craft-strategies-to-relax-lockdowns/
  2. Wei-jie Guan et al., Clinical Characteristics of Coronavirus Disease 2019 in China, New Engl. J. Med., Feb. 28, 2020, DOI: 10.1056/NEJMoa2002032
  3. David S. Jones, History in a Crisis — Lessons for Covid-19, New Engl. J. Med., Mar. 12, 2020, DOI: 10.1056/NEJMp2004361
  4. Tomas Pueyo, Coronavirus: Why You Must Act Now, Medium, Mar. 10, 2020, https://medium.com/@tomaspueyo/coronavirus-act-today-or-people-will-die-f4d3d9cd99ca
  5. Graziano Onder, Giovanni Rezza & Silvio Brusaferro, Case-Fatality Rate and Characteristics of Patients Dying in Relation to COVID-19 in Italy, JAMA, Mar. 23, 2020, doi:10.1001/jama.2020.4683
  6. John P.A. Ioannidis, A Fiasco in the Making? As the Coronavirus Pandemic Takes Hold, We Are Making Decisions Without Reliable Data, Stat, Mar. 17, 2020, https://www.statnews.com/2020/03/17/a-fiasco-in-the-making-as-the-coronavirus-pandemic-takes-hold-we-are-making-decisions-without-reliable-data/
  7. Eran Bendavid and Jay Bhattacharya, Is the Coronavirus as Deadly as They Say? Current Estimates About the Covid-19 Fatality Rate May Be Too High by Orders of Magnitude, Wall St. J., Mar. 24, 2020, https://www.wsj.com/articles/is-the-coronavirus-as-deadly-as-they-say-11585088464?mod=trending_now_3

Saturday, January 18, 2020

A Fourth Model of Law Enforcement Access to DNA Databases

First there were DNA databases of four or five RFLP VNTR measurements obtained by gel electrophoresis of the DNA fragments of convicted offenders. These were soon superseded by much larger databases for convicted-offenders (and later, arrestees) comprised of a larger number of STRs determined with capillary electrophoresis. The individuals supplying the DNA analyzed in these limited ways had no choice in the matter. Statutes compelled them to submit to DNA sampling, and only law enforcement authorities had access to these special-purpose, government-run databases.

In recent years, the potential for private databases to generate investigative leads through trawling (without individualized suspicion or probable cause) has begun to be exploited. Millions of SNP-based records reside on the servers of recreational genetics companies that process saliva samples with SNP arrays that determine the alleles present at hundreds of thousands of SNP loci. The customers of each direct-to-consumer (DTC) testing company can learn if other participating customers of that company have large blocks of DNA in common -- a situation indicative of a family relationship. (Other tests for relatedness in the private databases are also possible, but the haploblock-matching procedure has been the most productive for criminal investigations because it casts a wider net.) In addition, over a million samples reside in the voluntary private database known as GEDmatch, which enables genealogy enthusiasts to share their personal SNP-array data.

SNP-array testing does not work well with crime-scene DNA samples that contain mixtures of DNA from several individuals, that are severely limited in quantity, or that are degraded. However, an alternative technology, known as massively parallel, or next generation sequencing (MPS or NGS), can generate genomic data from challenging samples, and the sequence data can be culled to supply the SNP-array data that DTC companies would have provided had the source of the DNA evidence from the crime-scene sent a saliva sample in to one of these companies. Armed with such a data file, police who are able to access a DTC or GEDmatch-type database can look for potential relatives. By turning to these individuals (or public records about the families), police sometimes can find a suspect who merits further investigation.

Many DTC customers who did not realize their DNA might lead police to their relatives (known or unknown to them) have found this kind of forensic genomic genealogy (FGG) profoundly disturbing, at least when it was not something they had explicitly signed up for. Others have evinced less concern. The private database operators have responded with different policies. GEDmatch presumes that users object, but allows them to indicate their willingness to have their data used in law-enforcement trawls. Another possible response is an opt-out policy, in which the presumption is that individuals are not opposed to this use of their data. Yet another policy is that the database is always open to law enforcement, just as it is to curious individuals. Finally, the database could be completely closed to law-enforcement trawling (without judicial approval based on a showing of individualized cause).

In short, there are three main models of law enforcement access: (1) the government-run, law-enforcement-only database; (2) the private other-purpose, opt-in database; and (3) the private other-purpose, opt-out database. However, a fourth model is emerging -- a privately run law-enforcement-only database.

An article from the information service GenomeWeb reports on the plans of Othram, which bills itself as "the first technology company to apply all the power of modern sequencing and genomics to forensics" so as to secure "justice through genomics." The article includes a number of interesting statements from the company's CEO, David Mittelman. GenomeWeb explains that
Othram ... introduced DNASolves.com to solicit users of consumer genomics services to upload their data for the expressed desire to help law enforcement solve cold cases.

"Family Tree DNA is doing the opt-out model [with regards to law enforcement], GEDmatch is doing opt-in," said ... Mittelman. "I thought there should be another model," he said. "Since we do nothing but law enforcement, there is nothing to opt out of."

Mittelman, a former CSO at Family Tree DNA parent Gene by Gene ... said "I have enjoyed the consumer genetics and genealogy side, and I certainly enjoyed the medical side, but I saw the forensics market as a market that was underserved and could benefit from the technology that has been widely used and embedded in consumer and medical testing," he said. "It made sense to bring that technology over."

Mittelman credited the developments in the market with both the success of consumer genomics as well as advancements in next-generation sequencing technology. By some estimates, 30 million people have taken an array-based consumer test to date. Meantime, the drop in the price of sequencing, plus ongoing innovation in the field, means that it is now possible to perform whole-genome sequencing of highly degraded samples from crime scenes and then scour large databases to find genetic relatives, constructing genealogies to identify victims or perpetrators.
...
"In 2019, sequencing failed to penetrate the consumer market," said Mittelman. "But forensics is an interesting market where sequencing is superior to arrays," he said. "It is not just that sequencing gives you more information, in a lot of cases it's the only way to get information," he added.
Othram's approach is called Forensic Grade Genome Sequencing. According to Mittelman, array technology has a "high failure rate" when it comes to forensic samples, making sequencing the go-to technology when it comes to working with these kinds of samples.

"You really need special methods," said Mittelman. "We have developed proprietary methods to adapt the worst kinds of DNA to sequencing," he added. "I think in the long term forensics will only work with sequencing, while for consumer, arrays are good enough."

Othram has not yet published on its techniques, but eventually aims to do so, Mittelman said. While the company hones its sequencing capabilities, it is also hoping more customers of consumer services will be moved to upload their data to DNASolves.com. He noted that only a small percentage of those tested have elected to upload data to GEDmatch, meaning the potential exists to grow a new database of a different set of users interested in helping law enforcement.

"Rather than target a small number of people who are genealogy power users, and ask them to help solve crime instead, it seemed to me that you could approach the 30 million who have tested and tell them if you have tested and feel like getting involved, this is how you do it," said Mittelman. "You don't have to be a power user in genealogy to make a difference in crime solving."
Whether DNAsolves.com will attain the critical mass to be a useful investigative database is not guaranteed, so it is not clear that the data donors will be "making a difference in crime solving." Buut even if the database remains small, Othram can market its FGG service as offering the police agency access to a exclusive database. In a seeming excess of enthusiasm about finding "the most distantly related individuals," the website states that
We are all genetically related to one-another [sic]. Genetic genealogy uses DNA information in combination with genealogical and historical records to establish relationships between even the most distantly related individuals. This is a substantial improvement in human identification capability from current forensic testing methods that enable exact or near exact matches. When you contribute your DNA data, you help identify victims, missing persons, and perpetrators of crimes — even if you are a distant genetic relative.

REFERENCE

Justin Petrone, Forensic Genomics Market Advances Due to Consumer Databases, Technology Innovation, genomeweb, Jan. 9, 2020, https://www.genomeweb.com/sequencing/forensic-genomics-market-advances-due-consumer-databases-technology-innovation

Friday, January 17, 2020

What Are the Law Enforcement Implications of the GEDmatch Buyout?

GEDmatch is a free genetic genealogy database that allows people to upload their genomic data from direct-to-consumer (DTC) testing services such as 23andMe, Ancestry, MyHeritage, and Family Tree DNA. It permits cross-company searches for possible relatives among those who elect to participate. To date, roughly 1.3 million people have uploaded their SNP array data to the service, and GEDmatch continues to add about 1,000 people daily. 1/ Scientists and genealogists working for law enforcement agencies have been able to locate the source of DNA evidence from unsolved crimes by discovering a possible relative who supplied his or her data to public kinship searching in GEDmatch. The result has been fluctuating policies with respect to police use of GEDmatch's haploblock-matching software and uploaded data to produce some leads. 2/

Last month (19 Dec, 2019), GEDmatch's founder sent the following email to its registered users:
To GEDmatch users,

As you may know, on December 9 we shared the news that GEDmatch has been purchased by Verogen, Inc., a forensic genomics company whose focus is human ID. This sale took place only because I know it is a big step forward for GEDmatch, its users, and the genetic genealogical community. Since the announcement, there has been speculation about a number of things, much of it unfounded.

There has been concern that law enforcement will have greater access to GEDmatch user information. The opposite is true. Verogen has firmly and repeatedly stated that it will fight all unauthorized law enforcement use and any warrants that may be issued. This is a stronger position than GEDmatch was previously able to implement.

...It has been reported on social media that there is a mass exodus of kits from the GEDmatch database. There has been a temporary drop in the database size only because privacy policies in place in the various countries where our users reside require citizens to specifically approve the transfer of their data to Verogen. As users grant permission, that data will again be visible on the site. We are proactively reaching out to these users to encourage them to consent to the transfer.

... Verogen recognizes that law enforcement use of genetic genealogy is here to stay and is in a better position to prevent abuses and protect privacy than GEDmatch ever could have done on its own.

Bottom line: I am thrilled that the ideal company has purchased GEDmatch. The baby I created will now mature for the benefit of all involved. If anyone has any doubts, I may be reached at gedmatch@gmail.com. I will do my best to personally respond to all concerns.

Curtis Rogers
GEDmatch
Verogen's acquisition has produced angst among observers who worry that the company will exploit it for police purposes. The observation that Verogen "caters to law enforcement" 3/ has become a media meme. Verogen is a 2017 spin-off from Illumina, Inc., the San Diego Company that pioneered and acquisitioned its way to cheap DNA sequencing machinery for research and clinical applications. Illumina also sells the arrays that the DTC companies use, but that technology is not suitable for most crime-scene samples. Verogen's website proudly announces that "Verogen serves those who pursue the truth" by being "the world’s first sequencing company solely dedicated to forensic science. ... Powered by Illumina technology and free of legacy method allegiance, we are uniquely positioned to support forensic labs with innovative solutions purpose-built for the challenges of DNA identification."

The purpose-built solutions free of legacy method allegiance that Verogen now markets use Illumina's technology for massively parallel sequencing by synthesis that culminates in traditional autosomal forensic STR data and much more genomic data.It is superior to capillary electrophoresis, especially for small, degraded, and mixed DNA samples. Competitors market other packages of MPS devices, reagents, and software to forensic laboratories. 4/

How Verogen's ownership will affect police access to GEDmatch is not clear. No doubt, Verogen would like the police to use the database, but it cannot afford to alienate the genealogy enthusiasts who send in their SNP array data. It can provide materials and software that would help forensic laboratories who use sequencing technology to generate the SNP data in the format for trawling GEDmatch for haploblock matches. Indeed, Verogen has a "forensic genetic genealogy product in development for forensic laboratories [that] will actually provide more privacy protection for users of GEDmatch, as the test is focused on kinship analysis for forensic purposes." 5/ But I am curious as to how this purpose-built product will provide more privacy protection than occurs with kinship matching for private purposes.

NOTES
  1. Justin Petrone, Forensic Genomics Market Advances Due to Consumer Databases, genomeweb, Jan 9, 2020, https://www.genomeweb.com/sequencing/forensic-genomics-market-advances-due-consumer-databases-technology-innovation.
  2. Police Genetic Genealogy at GEDmatch: Is Opt-in the Best Policy?, Forensic Sci., Stat. & L., Sept. 21, 2019, https://for-sci-law.blogspot.com/2019/09/police-genetic-genealogy-at-gedmatch-is.html.
  3. Heather Murphy,  What You’re Unwrapping When You Get a DNA Test for Christmas, N.Y. Times, Dec. 22, 2019 ("The new owner, Verogen, said that it would actively fight future search warrants and that users can still opt out of helping police. But Verogen is also a company that has built its business, so far, on catering to law enforcement.").
  4. Brigitte Bruijns, Roald Tiggelaar & Han Gardeniers, Massively Parallel Sequencing Techniques for Forensics: A Review, 39 Electrophoresis 2642-2654 (2018), https://doi.org/10.1002/elps.201800082.
  5. Petrone, supra note 1.

Wednesday, January 15, 2020

OSAC-approved Testimony

According to the charter of the Department of Commerce's Organization of Scientific Area Committees for Forensic Science, better known as OSAC (but not to be confused with the Department of State's Overseas Security Advisory Council, also known as OSAC), "[t]he aims of the OSAC are to:
  • populate the OSAC Registry of Standards
  • promote the use of OSAC-endorsed standards by the forensic community, accreditation and certification bodies, and by the legal system
  • provide insight on each forensic science discipline’s research and measurement standard needs
  • enlist stakeholder involvement from a broad community; and to [sic]
  • establish and maintain working relationships with other similar organizations. 1/
OSAC is not a regulatory agency, and it does not review the work of forensic science laboratories or practitioners. Yet one court thought otherwise. In United States v. Lang, 2/ the defendant moved to exclude testimony from "the Government's firearms and toolmark expert—M.L. Cooper." The government wanted "Mr. Cooper ... to testify about what latent prints and DNA are and the factors that influence whether a latent print or DNA are left on a surface." Evidently, the point of the testimony was to explain the absence of fingerprint and DNA evidence in the case because of "the difficulty in obtaining latent prints and DNA from ammunition and the low rate at which such types of trace evidence are recovered."

Senior Judge Raymond L. Finch of the District Court of the Virgin Islands (formerly the court's Chief Judge) denied the motion. His unreported opinion on his pretrial ruling states that
On direct examination, Mr. Cooper testified that he reviewed a Crime Scene Evidence Report prepared by VIPD Officer Don Peter, photographs of the five buckshot shotgun shells recovered from Defendant's apartment, and photographs of the black garbage bag that contained the shells. He also testified that he confirmed his resulting opinion with OSAC (emphasis added).
This was not all that Mr. Cooper, and hence the court, had to say about OSAC. Judge Finch wrote that "Mr. Cooper is currently a member of the Association of Firearms and Toolmark Examiners—an organization that according to Mr. Cooper deals with the recovery of trace evidence—and maintains contact with the Organization of Scientific Area Committees ('OSAC')" (emphasis added). The nature of Mr. Cooper's "contact" with OSAC--or is it AFTE's "contact"--was left unstated, but the court thought it was important enough to add the footnote that "OSAC coordinates the 'development of standards and guidelines for the forensic science community to improve quality and consistency of work in the forensic science community.'"

Relying on an OSAC-approved standard for an opinion is one thing. But there are no standards on how to ascertain from a "review of the Crime Scene Evidence Report and photographs" that police would not be expected to recover a latent fingerprint or an adequate quantity of DNA to analyze, and being in "contact with" OSAC does not tell the court anything about the expertise of a witness or an organization.

Other parts of the opinion are more comprehensible (but not necessarily correct). The court wrote that the testimony "satisfies the reliability test under Daubert" just because it "will be based upon his more than thirty years of experience working in the field of forensic science." Experience can be a good thing, but it is not the scientific validity discussed in Daubert. More plausibly, the court moved away from its equation of experience to science. It stated that Mr. Cooper's methodology did not have to satisfy Daubert after all. "Rather than [applying] a methodology that satisfies the requirements of Daubert," Mr. Cooper would do little more than recount his personal experience with "the recovery of trace evidence" That much, the court insisted, "is entirely appropriate under Rule 702. See United States v. McNeil, 2010 U.S. Dist. LEXIS 290, at *8, 2010 WL 834667 (M.D. Pa. Jan. 5, 2010) ('To the extent that the expert has knowledge of the frequency of firearms without latent prints, the expert [can] testify to that knowledge.')."

NOTES
  1. OSAC Charter and Bylaws, Sept. 26, 2019, Version 1.6.
  2. Crim. Action No. 2015-0013, 2016 WL 1734087 (D. V. I., Apr. 28, 2016).