Sunday, October 25, 2015

SWGDAM Guidelines on "Probabilistic Genotyping Systems" (Part 2)

What makes a "Probabilistic Genotyping System" probabilistic? That a computer program delivers a probability related to a DNA profile does not make it a PGS. After all, traditional, manual analysis of DNA data leads to probabilities. Here, I present a toy example of a single-source sample to convey a sense of the nature of probabilistic genotyping.

I do so with some trepidation. Neither the SWGDAM Guidelines nor the articles that I have located supply a simple and clear exposition of the actual workings of any modern forensic PGS. The Guidelines state that
A probabilistic genotyping system is comprised of software, or software and hardware, with analytical and statistical functions that entail complex formulae and algorithms. Particularly useful for low-level DNA samples (i.e., those in which the quantity of DNA for individuals is such that stochastic effects may be observed) and complex mixtures (i.e., multi-contributor samples, particularly those exhibiting allele sharing and/or stochastic effects), probabilistic genotyping approaches can reduce subjectivity in the analysis of DNA typing results.
That sounds great, but what do these "complex formulae and algorithms" do? Well,
probabilistic approaches provide a statistical weighting to the different genotype combinations. Probabilistic genotyping does not utilize a stochastic threshold. Instead, it incorporates a probability of alleles dropping out or in. In making use of more genotyping information when performing statistical calculations and evaluating potential DNA contributors, probabilistic genotyping enhances the ability to distinguish true contributors and noncontributors.
Moreover, "[t]he use of a likelihood ratio as a reporting statistic for probabilistic genotyping differs substantially from binary statistics such as the combined probability of exclusion."

This sounds good too, but what is "a statistical weighting," and how is a probability of exclusion, which is not confined to 0 to 1, a "binary statistic"? To gain a clearer picture of what might be going on, I thought I would start with the simplest possible situation — a crime-scene sample with a single contributor — to surmise how a probabilistic analysis might operate. My analysis is something of a guess. Corrections are welcome.

Two Peaks, One Inferred Genotype, One Likelihood Ratio of 50: Not a PGS!

In "short tandem repeat" typing via capillary electrophoresis, the laboratory extracts DNA from a sample and uses the PCR (polymerase chain reaction) to make millions of copies of a short stretch of DNA between a designated starting point and a stopping point (a "locus"). These fragments vary in length among different individuals (although none are unique). The laboratory runs the sample fragments through a machine that measures the quantity of the fragments as a function of the length of the fragments. For example, a plot of the quantity on the y-axis and the fragment length on the x-axis might show two prominent peaks, which I will call A and B, of roughly equal height rising above a noisy baseline. This AB pattern at a single locus is exactly what one would expect for DNA from an individual who inherited a fragment of length A from one parent and a fragment of length B from the other parent. Starting with roughly equal numbers of maternally and paternally inherited DNA molecules in the original sample, PCR should generate about equal quantities of the maternal and paternal length variants ("STR alleles") of the two distinct lengths. These produce the two peaks in the graph (the electropherogram).

The analyst then could compute the “random match probability” or “probability of inclusion” (PI) — that is, the probability P(RAB) that a randomly selected individual would be type AB. Even if the analyst used a computer program to do the calculation, no “probabilistic genotyping” would be involved. The “genotype” AB would be regarded as known to a certainty (for the purpose of the computation), and the probability PI pertains to something else — to the chance of coincidentally finding an individual with a matching profile: PI = P(RAB). If 1 in 50 people have the profile AB, then PI = 1/50.

The evidentiary value of the inclusion can be computed as a “likelihood ratio” (LR). If the hypothesis (Hp) that the suspect, who also is type AB, is the contributor of the DNA in the sample is correct, and if the sample has plenty of undegraded DNA, the probability of the data DAB (an A and a B peak detected in the sample) is P(DAB|Hp) = 1. On the other hand, if someone unrelated to the suspect is the contributor (Hd), then P(DAB|Hd) is the probability of inclusion PI = 1/50. Thus, the evidence — the A and B peaks — is 1/PI = 50 times more probable when the suspect is the contributor than when an unrelated person is. This ratio of the probabilities of the evidence conditional on the hypotheses is the likelihood ratio. It measures the support the evidence lends to Hp as opposed to Hd. LRs greater than 1 support Hp over Hd (e.g., Kaye et al. 2011).

Two Peaks, Two Inferred Genotypes with Probabilities for Each Genotype: A PGS?

This much is straightforward, conventional thinking. But an AB contributor is not the only conceivable explanation for the two peaks. Maybe they reflect DNA from an AA individual (one who inherited the fragment of length A from both parents), and the B is just an artifact known as “stutter” (Brooks et al. 2012). If this possibility cannot be dismissed as wildly improbable (as it could be if, for example, the putative stutter peak were far from the A peak), then the analysis should take into account both AA and AB as possible contributor profiles.

One way to do so would be to study the detection probability P(DAB) in experiments with samples from AA and AB contributors. Suppose that a large number of such experiments showed that when the contributor is AA, the probability of detecting AB is P(DAB|CAA) = 1/10 and that when the contributor is AB, the probability is P(DAB|CAB) = 1. Sometimes, AA contributors produce AB peaks; AB contributors always do.

In a case in which the suspect is type AB, what is the evidentiary value of the two peaks A and B? The suspect is still AB, so P(DAB|Hp) is unchanged at 1. But the denominator of the LR, P(DAB|Hd) requires us to consider the probability that the contributor’s profile is AA as well as the probability that it is AB. Imagine that the laboratory receives crime-scene samples with DNA profiles that are representative of a population in which 1 in 100 people are AA and (as stated before) 1 in 50 are AB. Because only 1 in 10 DNA samples from AA contributors will appear to be AB, about 1 in 1000 samples will have the AB peaks and come from AA contributors:

P(CAA & DAB) = P(CAA) ⋅ P(DAB|CAA) = (1/100) ⋅ (1/10) = 1/1000.

More samples, about 20 per 1000, will have the AB peaks and come from AB contributors:

P(CAB & DAB) = P(CAB) ⋅ P(DAB|CAB) = (1/50) ⋅ (1) = 20/1000.

Thus, in about 20 out of 21 detections of AB peaks, the contributor is AB. (Most readers who have borne with me this far will recognize this result as a simple application of Bayes' rule for the posterior probability: P(CAB|DAB) = 20/21.)

A PGS thus could assign probabilities of P(CAA|DAB) = 1/21 and P(CAB|DAB) = 20/21 for the two possible contributor genotypes. The hypothesis Hd is that either an unrelated person who is AA or, as before, that the peaks come from an unrelated AB contributor. If the suspect is not the source and if the apparent AB profile really is AA (which has probability 1/21), Hd requires that a random, unrelated person be type AA (an event that has probability P(RAA) = 1/100). Likewise, if the suspect is not the source and the apparent AB profile really is AB (which has probability 20/21), then Hd requires that a random, unrelated person be type AB (an event that has probability P(RAB) = 1/50). Consequently, the probability of the evidence DAB given Hd is

         P(DAB|Hd) = P(RAA) ⋅ P(CAA|DAB) + P(RAB) ⋅ P(CAB|DAB)
                    = (1/100) (1/21) + (1/50) (20/21) = 41/2100 = 0.0195.

This likelihood is very close to the previous denominator of 1/50 = 0.020. The resulting LR is 2100/41 = 51.2.

The Probability in PGS

This toy model of a PGS only used information about peak location and only mentioned a stutter peak as a source of uncertainty in the contributor's genotype. A more sophisticated PGS would use peak heights as well and would attend to allelle drop-in and drop-out, and other complicating features. The most complete models dispense with the rules of thumb (“analytical thresholds,” “stochastic thresholds,” and “peak-height ratios”) that human examiners employ to decide whether a peak is high enough to count as real, what to do with it in computing a likelihood ratio, and what potential genotypes to cross off the list of possibilities when confronted with a mixture of DNA from several contributors (Kelly et al. 2014).

I do not propose to explain these matters any better than SWGDAM has. My purpose here has been to clarify just what is “probabilistic” about a PGS. The key point is not that the system produces a likelihood ratio as opposed to a probability of exclusion or inclusion. Likelihood ratios also apply to categorical inferences as to what profiles are present in a mixed sample. A PGS is distinctive because it assigns probabilities to the possible profiles and uses more information to arrive at what, one hopes, is a better likelihood ratio for the hypotheses about whether a suspect is a contributor.

References
  • C. Brookes, J.A. Bright, S. Harbison, J. Buckleton, Characterising Stutter in Forensic STR Multiplexes, 6 Forensic Sci. Int’l: Genetics 58-63 (2012)
  • David H. Kaye et al., The New Wigmore on Evidence: Expert Evidence (2d ed. 2011)
  • Hannah Kelly, Jo-Anne Bright, John S. Buckleton, James M. Curran, A Comparison of Statistical Models for the Analysis of Complex Forensic DNA Profiles, 54 Sci. & Justice 66–70 (2014)
Acknowledgement
Thanks are owed to Sandy Zabell for correcting errors in the original posting. This version was last updated 1 February 2016.

Thursday, October 22, 2015

SWGDAM Guidelines on "Probabilistic Genotyping Systems" (Part 1)

In June, the Scientific Working Group on DNA Analysis Methods (SWGDAM), approved new “Guidelines for the Validation of Probabilistic Genotyping Systems.” 1/ They begin,
Guidance is provided herein for the validation of probabilistic genotyping software used for the analysis of autosomal short tandem repeat (STR) typing results. These guidelines are not intended to be applied retroactively. It is anticipated that they will evolve with future developments in probabilistic genotyping systems.
These three sentences, raise four questions. First, is the phrase “probabilistic genotyping system” (PGS) the best label? I will get to the question of what “probabilistic” means a little later, but given the perception of segments of the public and the legal community that “autosomal short tandem repeat (STR) results” are “very likely” “to reveal predispositions to diseases in the individuals being profiled as well as their siblings and offspring,” 2/ is “genotyping” the right word to use for identifying DNA variations that are not genes? A more neutral term such as “probabilistic typing systems” might be less suggestive.

Second, why do the drafters of standards and guidelines prefer stilted writing—“guidance is provided herein”—as opposed to plain English sentences such as “This document offers guidance”? I know this kind of criticism is small potatoes, but scientists are smart enough to be good writers.

Third, what are the drafters trying to say with the doubly passively voiced sentence, “These guidelines are not intended to be applied retroactively”? Who should not apply these standards retroactively? One would think that the guidelines are for laboratories, but how could a laboratory apply a recommendation retroactively? It cannot go back in time to validate software that it has been using even though neither it nor the developer had validated the software in the manner that SWGDAM now recommends. The only thing the laboratory could do to give retroactive effect to the new advice would be to use some better validated software on data from old cases and advise prosecutors, defendants, or defense lawyers of major discrepancies. Is SWGDAM saying that looking back at past cases (for research or other purposes) would be wrong? Or merely that SWGDAM is taking no position on the desirability of undertaking such retrospective analyses? Or is this part of the guidelines written for a difference audience—courts that might be asked to grant postconviction relief? But unless every PGS was adequately validated, surely courts should consider what these guidelines have to say as relevant to (but not necessarily dispositive of) whether the laboratory’s earlier report was scientifically acceptable. Most courts can be expected to appreciate the fallacy of the argument that "because the world gets wiser as it gets older, therefore it was foolish before." 3/

Fourth, why does SWGDAM anticipate that “future developments in probabilistic genotyping systems” will cause these standards to “evolve”? The principles of good software development and validation do not depend on the specific programs. Those principles may evolve whether or not PGSs improve over time. Of course, the guidelines could change if the programs become so superior that SWGDAM would reconsider its view (expressed in the next paragraph) that the only permissible use of a PGS is “to assist the DNA analyst in the interpretation of forensic DNA typing results.” Is SWGDAM envisioning that it could reverse its opinion that “Probabilistic genotyping is not intended to replace the human evaluation of the forensic DNA typing results” because of “future developments in [PGS]”? In light of current problems with human interpretations of mixtures of minute quantities, there are observers who would welcome replacing the current protocols for interpreting these samples with valid and reliable automated expert or probabilistic systems.

Notes

1. Scientific Working Group on DNA Analysis Methods, Guidelines for the Validation of Probabilistic Genotyping Systems, June 15, 2015

2. Gary R. Skusea1 & Anne M. Burgera, Justice as Fairness: Forensic Implications of DNA and Privacy, Champion, Apr. 2015, at 24. For a more authoritative assessment, see Henry T. Greely & David H. Kaye, A Brief of Genetics, Genomics and Forensic Science Researchers in Maryland v. King, 53 Jurimetrics J. 43 (2013).

3. Hart v. Lancashire &Yorkshire Ry. Co., 21 L.T.R. N.S. 261, 263 (1869).

Saturday, August 22, 2015

Disentangling Two Issues in the Hair Evidence Debacle

Forensic-science practitioners commonly present findings regarding traces left at crime-scenes, on victims or suspects, or on or in their possessions. Such trace evidence can take many forms. Physical traces such as fingerprints, striations on bullets, shoe and tire prints, and handwritten documents are common examples. Biological materials, such as blood, semen, saliva, and hairs also are fodder for the crime laboratory. Comparisons of a questioned and known sample can supply valuable information on whether a specific suspect is associated in some manner with a crime. Viewers of the acronymious police procedurals—NCIS, CSI, and Law and Order SVU—know all this.

For decades, however, legal and other academics have questioned the hoary courtroom claims of absolutely certain identification of one and only one possible source of trace evidence. In the turbulent wake of the 2009 report of a National Research Council committee, these views have slowly gained traction in the forensic-science community. Indeed in the popular press and among investigative reporters, the pendulum may be swinging in favor of uncritical rejection of once unquestioned forensic sciences. Recent months have seen an episode from Frontline presenting DNA evidence as "anything but proven"; 1/ they have included unfounded reports that as many as 15 percent of men and women imprisoned with the help of DNA evidence at trial are wrongfully convicted; 2/ and award-winning journalists have spread the word that the FBI "faked an entire field of forensic science," 3/ placed "pseudoscience in the witness box," 4/ and palmed off "virtually worthless" evidence as scientific truth. 5/

The last set of reports stem from an ongoing review of well over 20,000 cases in which the FBI laboratory issued reports on hair associations. The review spans decades of hair comparisons, and it is showing so many questionable statements that the expert evidence on hair associations stands out as "one of the country's largest forensic scandals." 6/ Its preliminary findings provoked prominent Senators to speak of an "appalling and chilling ... indictment of our criminal justice system" 7/ and to call for a "root cause analysis" of ubiquitous errors. 8/ A distressed Department of Justice and FBI joined with the Innocence Project and the National Association of Defense Lawyers not only to publicize these failings, but also to call on states "to conduct their own independent reviews where ... examiners were trained by the FBI." 9/ Projecting the outcome of cases that have yet to be reviewed, postconviction petitions refer ominously to "[t]housands of . . . cases the Justice Department now recognizes were infected by false expert hair analysis" 10/ and "pseudoscientific nonsense." 11/

The hair scandal illustrates two related problems with many types of forensic-science testimony. The first is the problem of foundation—What reasons are there to believe that hair or other analysts possess sufficient expertise to produce relevant evidence of associations between known and unknown samples? For years, commentators and some defense counsel have posed legitimate questions (with little impact in the courts) about the reliability and validity of physical comparisons by examiners asked to judge whether known and unknown samples are similar in enough respects—and not too dissimilar in other respects—to support a claim that they could have originated from the same individual. To paraphrase Gertrude Stein, is there enough there there to warrant any form of testimony about a positive association? This is the existential question of whether, in the words in of the Court in Daubert v. Merrell Dow Pharmaceuticals, "[t]he subject of an expert's testimony [is] 'scientific . . . knowledge.'" 12/ Or, at the other extreme, is the entire enterprise ersatz—a "fake science" and a "worthless" endeavor?

As I see it, the harsh view that physical hair comparisons are pure pseudoscience, like astrology, graphology, homeopathy, or metoposcopy, is not supportable. The FBI's review project itself rests on the premise that hair evidence has some value. If the comparisons were worthless—like consulting the configurations of the stars or reading Tarot cards—there would be no need to review individual cases. In all cases of an association, the FBI would have exceeded the limits of science. But this only shows that the FBI thinks that there is a basis for some testimony of a positive association. What evidence supports this belief?

One reason to think that the FBI's hair analysts generally possess some expertise in making associations comes from an intriguing study (usually cited as proof of the failings of microscopic hair comparisons) done more than ten years ago. 13/ FBI researchers took human hairs submitted to the FBI laboratory for analysis between 1996 and 2000 and mitotyped them (a form of DNA testing) whenever possible. The probability of an FBI examiner finding of a positive association when comparing hairs from the same individual (as shown by mitotyping) exceeded the probability of this finding when comparing hairs from different individuals by a factor of 2.9 (with a 95% confidence interval of 1.7 to 4.9). A test with this performance level only supplies evidence that is, on average, weakly diagnostic of an association. Still, it is not simply invalid. 

The second problem lies in the presentation of perceived associations. Even if there is a there there, are forensic-science practitioners staying within the boundaries of their demonstrated expertise? Or are they purporting to know more than they do? The FBI's revelations about hair evidence are confined to this issue of overclaiming. What the FBI has uncovered are expert assertions in one case after another that are said to outstrip the core of demonstrated knowledge. Such overclaiming is one form of scientifically invalid testimony, 14/ but it is not the equivalent of an entire invalid science.

In considering the pervasiveness of the problem of overclaiming, the FBI's figure of 90+ percent is startling. It is so startling that one should ask whether it accurately estimates the prevalence of scientifically indefensible testimony. There are reasons to suspect that it might not. After all, the Hair Comparison Review Project was not designed to estimate the proportion of cases in which FBI examiners gave testimony that was, on balance, scientifically invalid. It is intended to spot isolated statements that claimed more than an association that, for all we know, could be quite common in the general population. The general criteria for judging whether particular testimony falls into this category have been publicized, but the protocol and specific criteria that the FBI reviewers are using have not been revealed. No representative sample of the reports or transcripts that are judged to be problematic is available, but a few transcripts in the cases in which the FBI has confessed scientific error indicate that at least some classifications are open to serious question.

It may be instructive to contrast the response to two very similar statements noted in a posting of May 23, 2015, about the court-martial of Jeffrey MacDonald: (1) "this hair ... microscopically matched the head hairs of Colette MacDonald"; and (2) "[a] forcibly removed Caucasian head hair ... exhibits the same microscopic characteristics as hairs in the K2 specimen. Accordingly, this hair is consistent with having originated from Kimberly MacDonald, the identified source of the K2 specimen." The first statement passed muster. The second did not.

If the review in this case is not aberrant, examiners can say that two hairs share the same features—they "match" and are consistent with one another—but they must not add the obvious (and scientifically undeniable) fact that this observation (if correct) means that they could have had the same origin or that they are "consistent with" this possibility.

Of course, one can criticize phrases like "consistent with" and "match" as creating an unacceptable risk that (in the absence of clarification on direct examination, cross-examination, or by judicial instruction) jurors will think the words connote a source attribution. But arguments of this sort stray from determinations that an examiner has made statements that "exceed the limits of science" (the phrase the Justice Department uses in confessing overclaiming). They represent judgments that an examiner has made statements that are scientifically acceptable but prone to being misunderstood.

To be sure, this latter danger is important to the law. It should inform rulings of admissibility under Rules of Evidence 403 and 702. It is a reason to regulate the manner in which experts testify to scientifically acceptable findings, as some courts have done. Laboratories themselves should adopt and enforce policies to ensure that reports and testimony avoid terminology that is known to convey the wrong impression. But it is misleading to include scientifically acceptable but psychologically dangerous phrasing in the counts of scientifically erroneous statements. Case-review projects ought to flag all instances in which examiners have not presented their findings as they should have, but reports ought to differentiate between statements that directly "exceed the limits of science" and those that risk being misconstrued in a way that would make them "exceed the limits of science." One size does not fit all.

NOTES

This posting is an abridged and modified version of a forthcoming essay about the FBI's Microscopic Hair Review Comparison Review. A full draft of the preliminary version that is being edited for publication is available. Comments and corrections are welcome, especially before publication while there is time to improve the essay.

  1. Some inaccuracies in the documentary are noted in a June 24, 2015, posting on this blog.
  2. See the posting of July 3, 2015 on the blog (debunking this initial assertion of the Rand Corporation).
  3. Dahlia Lithwick, Pseudoscience in the Witness Box: The FBI Faked an Entire Field of Forensic Science, Slate (Apr. 22, 2015 5:09 PM).
  4. Id. These "shameful, horrifying errors" comprised "a story so horrifying . . . that it would stop your breath." Id.
  5. Erin Blakemore, FBI Admits Pseudoscientific Hair Analysis Used in Hundreds of Cases: Nearly 3,000 Cases Included Testimony About Hair Matches, a Technique that Has Been Debunked, Smartnews (Apr. 22, 2015) (quoting Ed Pilkington, Thirty Years In Jail For A Single Hair: The FBI's 'Mass Disaster' of False Conviction, Guardian, Apr. 21, 2015.
  6. Spencer S. Hsu, FBI Admits Flaws in Hair Analysis over Decades, Wash. Post, Apr. 18, 2015.
  7. Id. (quoting "Sen. Richard Blumenthal (D-Conn.), a former prosecutor").
  8. Spencer S. Hsu, FBI Overstated Forensic Hair Matches in Nearly All Trials Before 2000, Wash. Post, Apr. 19, 2015 (quoting "Senate Judiciary Committee Chairman Charles E. Grassley (R-Iowa) and the panel's ranking Democrat, Patrick J. Leahy (Vt.)").
  9. FBI, FBI Testimony on Microscopic Hair Analysis Contained Errors in at Least 90 Percent of Cases in Ongoing Review (Apr. 20, 2015).
  10. Petition for a Writ of Certiorari, at 2-3, Ferguson v. Steele, 134 S.Ct. 1581 (2014) (No. 13-1069).
  11. Id. at 19. Justice Breyer saw the “errors” referred to in the press release as emblematic of “flawed forensic testimony” generally and a reason to hold the death penalty unconstitutional. Glossip v. Gross, No. 14-7955 (June 29, 2015) (dissenting opinion).
  12. 509 U.S. 579, 590 (1993).
  13. Max M. Houck & Bruce Budowle, Correlation of Microscopic and Mitochondrial DNA Hair Comparisons, 47 J. Forensic Sci. 1 (2002).
  14. For this vocabulary, see Brandon Garrett & Peter Neufeld, Invalid Forensic Science Testimony and Wrongful Convictions, 95 Va. L. Rev. 1 (2009); cf. Eric S. Lander, Fix the Flaws in Forensic Science, N.Y. Times, Apr. 21, 2015 (“The F.B.I. stunned the legal community on Monday with its acknowledgment that testimony by its forensic scientists about hair identification was scientifically indefensible in nearly every one of more than 250 cases reviewed.”).

Monday, August 17, 2015

First NIST OSAC Forensic Science Standards Up for Public Comment

One of the responses to the 2009 NRC report on forensic science was the creation last year of an Organization of Scientific Area Committees (OSAC) for forensic science organized by the National Institute of Standards and Technology (NIST). This organization is developing new standards for forensic disciplines to follow.

Last week, NIST opened a 30-day public comment period for five standards from the Chemistry Scientific Area Committee. They are continuations or updates of existing ASTM (American Society of Testing and Materials) standards. The NIST OSAC News Release on the public comment period is at http://www.nist.gov/forensics/osac/osac-opens-public-comment.cfm. The five standards under consideration for inclusion on the OSAC Registry of Approved Standards are as follows:
  • ASTM E2329-14 Standard Practice for Identification of Seized Drugs
  • ASTM E2330-12 Standard Test Method for Determination of Concentrations of Elements in Glass Samples Using Inductively Coupled Plasma Mass Spectrometry (ICP-MS) for Forensic Comparisons
  • ASTM E2548-11e1 Standard Guide for Sampling Seized Drugs for Qualitative and Quantitative Analysis
  • ASTM E2881-13e1 Standard Test Method for Extraction and Derivatization of Vegetable Oils and Fats from Fire Debris and Liquid Samples with Analysis by Gas Chromatography-Mass Spectrometry
  • ASTM E2926-13 Standard Test Method for Forensic Comparison of Glass Using Micro X-ray Fluorescence (ยต-XRF) Spectrometry
Although they may seem technical and have forbidding names, some of these proposed standards should be of interest to lawyers as well as forensic scientists and statisticians who might want them to address how findings should be presented in court or in reports. For example, one standard involving glass fragments requires a difference of 3 standard deviations before the analyst can reject the hypothesis that the fragments on the suspect came from the crime scene. But 2.9 usually would be pretty good evidence that the suspect's fragments are from some other glass. What should the standard require or allow an expert to report in cases like this, where the suspect's fragments lie within the broad window for measurement error? Should there be an adjustment to the rejection range if more than one fragment has been tested? Should there even be a fixed window, or should the analyst simply report the probability of differences in the measurements as or more extreme as those observed if the fragments on the suspect came from the crime-scene glass? Better still, can a likelihood ratio be provided?

It appears that this 30-day period also offers an opportunity to view related ASTM standards on forensic science tests. Normally, ASTM, as the copyright holder, does not make its standards freely available.

Directions for subscribing to the OSAC newsletter and receiving announcements of comment periods, new standards, etc., are at the above URL and at http://www.nist.gov/forensics/osac/osac-launches-monthly-newsletter.cfm.

Disclosure and disclaimer: Although I am a member of the Legal Resource Committee of OSAC, the views expressed here (to the extent I have expressed any) are mine alone. They are not those of any organization. They are not necessarily shared by anyone inside (or outside) of NIST, OSAC, any SAC, any OSAC Task Force, or anyone else in the Legal Resource Committee.

Friday, July 24, 2015

What Proves that "the expert and his methods couldn’t possibly be reliable"?

I am developing an allergic reaction to the following kind of argument: "A forensic-science expert testified that a trace at the crime scene or on the victim was associated (to some degree of certainty) with the defendant. DNA later evidence exonerated the defendant. Therefore, the expert’s methods couldn’t possibly be reliable."

This reasoning is not very different from saying that a pitcher who does not strike out every batter couldn’t possibly be a reliable pitcher; that a polling firm that fails to correctly predict every election must be using methods that couldn’t possibly be reliable; or that a test for heart disease that sometimes errs couldn’t possibly be reliable. Without considering the success as well as the failure rate of the method, it is impossible to say that it is unreliable—or, by the same token, that it is reliable (in the sense of being worth relying on). 1/

Yet, the argument from cases of exonerations—lacking any comparison group—are legion in discourse on the use of trace evidence for identification. The latest example I encountered comes from a Washington Post blog site. Two days ago, Radley Balko wrote that
[A] defendant was convicted due to the testimony of a forensic expert who claimed that his “science” showed the defendant, and only the defendant, could have committed the crime. That conviction was later upheld by an appeals court in an opinion that explained in detail why the expert and his methods were legitimate and reliable. The defendant was later exonerated by DNA testing, thus demonstrating that the expert and his methods couldn’t possibly be reliable.
Mr. Balko went on to write that “not only did the courts continue to allow bite mark matching into evidence, every single time a defendant challenged its validity, that defendant lost.” 2/

Both the reasoning and the description of legal history are not quite right. To begin with, even absolute proof of innocence only shows that the test has a nonzero false-positive error rate (no surprise there) and that the witness should not have claimed to a certainty that no one else could have left the mark in question.

To be sure, some evidence that has found breathing space in the courtroom should be squeezed out entirely. But let’s face it—no scientific test meets the standard of perfection. If every case in which evidence that has produced false convictions meant “that the expert and his methods couldn’t possibly be reliable,” there could be no evidence. Overselling has occurred with every type of forensic evidence—from bitemarks to toolmarks to fingerprints to DNA. Courts, scientists, and criminalists should do their best to prevent this. Thus, whether through rules of evidence or through education and monitoring of analysts, testimony must be calibrated to the power of the scientific technique. The testimony should fairly express the known probative value of evidence from a validated method.

An example of testimony that violates this precept comes from Ege v. Yukins. 3/ In that case, a discredited dental expert testified as follows:
Q: Now, Doctor, with regard to your testimony, you indicated that it's highly consistent with the dentition of Defendant Carol Ege; is that correct?
A: Yes.
Q: Okay. With regard to—let me ask you a question. Let's say you have the Detroit Metropolitan Area, three, three and a half million people. Would anybody else within that kind of number match like she did?
A: No, in my expert opinion, nobody else would match up. 4/
The eventual outcome in the case contradicts the assertion that no challenge to the admission of bitemark evidence has succeeded. The state trial judge in Ege realized that the testimony was improper and only “denied [postconviction] relief because of the lack of a contemporaneous objection and a view that the showing of prejudice was insufficient.” 5/

A federal district court also concluded that “expert testimony identifying the petitioner as the only possible perpetrator of the alleged bite mark in the Detroit metropolitan area was improperly admitted.” 6/

The U.S. Court of Appeals for the Sixth Circuit agreed “with the district court that ‘Dr. Warnick's opinion that the petitioner was the only person in the entire Detroit metropolitan area who could have made the mark on the corpse carried an aura of mathematical precision pointing overwhelmingly to the statistical probability of guilt, when the evidence deserved no such credence.’” 7/ It affirmed the order for a new trial.

In short, the argument that a method of forensic identification that has been proved to be fallible is, for that reason alone, inadmissible proves too much. Likewise, the claim that no challenge to bitemark evidence has ever prevailed is exaggerated (although not by much). 8/

Please do not misunderstand me. The series of articles on bitemark evidence from which the remarks I have quoted were taken is impressive and useful. In offering these corrections to two small parts that seem a bit extreme, I am not arguing that bitemark analysis, which has little claim to validity, is either reliable (in the statistical sense that repeated analyses of the same marks give the same answers) or valid (in the sense that the answers are more often correct when marks from the same source are analyzed than when marks from different sources are compared). From the writing I have seen, bitemark analysis does not cut it.

I also believe that cases of false convictions should be studied and that the existence of a given type of scientific evidence in these cases should not be ignored. Finding a large number of false convictions with such evidence present is a warning signal. The evidence may come from a method that has a large false-positive rate, 9/ and that possibility must be investigated to decide whether the evidence should be excluded across the board or whether juries should receive the information -- together with an honest and clear explanation of the uncertainty in the results.

NOTES
  1. See infra note 9.
  2. Radley Balko, A High-ranking Obama Official Just Called for the “Eradication” of Bite Mark Evidence, The Watch, Wash. Post, July 22, 2015.
  3. 485 F.3d 364 (6th Cir. 2007).
  4. Ege v. Yukins, 380 F.Supp. 2d 852, 871 (E.D. Mich. 2005), affirmed in part, reversed in part, 485 F.3d 364 (6th Cir. 2007).
  5. Id. at 857–58.
  6. Id. at 858
  7. 485 F.3d at 376.
  8. The federal courts in Ege treated the answer to the 3.5 million people as "probability testimony" without questioning Michigan's general rule that bitemark identifications are admissible. A true (and wrongly decided) case of bitemark probability evidence is State v. Garrison, 585 P.2d 563 (Ariz. 1978).
  9. The false-positive probability is P(+|O), where + is a positive statement ("the defendant left the mark") and O is the fact that some other person left the mark. Even if this probability is small, a disturbing number of false convictions could involve this evidence. Suppose that P(+|O) = 0.02, that 1,000 tests are performed in a set of cases with marks, and guilty defendants left the marks in 60% of these cases. The expected number of false positives is (0.02)(400) = 8. Assume that the probability of a true positive is P(+|S) = 0.96, where S means that the defendant is the source of the mark. Then the expected number of true positives is (600)(0.96) = 576. If defendants are convicted in all these cases, 8 convictions will be false (assuming that the culprit left the mark), and the many true positives will not be seen in the cases of exonerations of the innocent defendants. As indicated at the outset of these remarks, other data than exonerations are required to judge whether the test is reliable and valid.