Showing posts with label cognitive bias. Show all posts
Showing posts with label cognitive bias. Show all posts

Saturday, January 15, 2022

Bones of Contention: A Standard for Analyzing Skeletal Trauma in Forensic Anthropology

The Academy Standards Board (ASB) of the American Association of Forensic Sciences (AAFS) posted the second proposed draft of a "Standard for Analyzing Skeletal Trauma in Forensic Anthropology" for public comment. The standard does not go far toward standardizing procedures or showing that the procedures to which it applies have been scientifically tested. Of course, it could well be that ample, well designed studies have demonstrated that forensic anthropologists can consistently and accurately classify skeletal defects in human remains according to the categories the standard mentions. But the standard contains no bibliography and no citations to show that this is the case. 

It contains some negative injunctions and a few positive suggestions about reporting -- for example:

  • Forensic anthropologists shall not determine cause or manner of death.
  • Practitioners shall not estimate the temperature or duration of heat exposure based on thermal defects to bone.
  • Practitioners may report the minimum number of traumatic events (e.g., blunt impacts, projectile entry defects, or sharp defects) observed skeletally, but shall not report a definitive maximum number of impacts, as skeletal trauma evidence may not reflect all impacts to the body.
  • When a suspect tool is submitted for analysis, similarities between the tool and defect may be reported; conclusions shall be reported in terms of an exclusion or failure to exclude.

As such, ASB 147-21 is not without any redeeming legal value. Nevertheless, it does not articulate any analytical process by which the classifications it calls for should be made (cf. "vacuous standards"); it requires no reporting of the uncertainty in this process; it does not contemplate the possibility of evidence-based rather than conclusion-based statements of the implications of the data; and it refers to an all-inclusive list of methods as "acceptable." If I may elaborate:

File:Human skeleton remains.jpg - Wikimedia Commons

Is "Interpretation" Limited to an Opinion on the Inference (Conclusion) from the Data?

The revision defines "trauma interpretation" as "Opinion regarding the mechanism of, timing, direction of impact(s) or minimum number of impacts associated with skeletal defect(s) based on quantitative and/or qualitative observations." The phrase "based on ... observations" indicates that the opinion expresses a belief in the truth, falsity, or probability of an inference being drawn from the data. Interpretation should include the possibility of describing the strength of the evidence in favor of the inference rather than opining on the truth, falsity, or probability of the conclusion itself. In addition, if the opinion-statement is an assertion that the hypothesis about what happened is true or false (either categorically or to some probability), it is not just based on the data but on a prior probability for the hypothesis as well.

Despite these definitions, the standard sanctions "interpretation" in the form of rudimentary statements about the extent to which the data prove the hypothesis in question. Section 6 notes that "Trauma interpretation shall be clearly identified in the report using terms such as ‘indicative of’ and ‘consistent with’ or by using a subheading titled ‘Interpretation.’"These phrases have their problems, but they are one manner of referring to the probability of the evidence given the truth of certain probabilities rather than vice versa.

Is Interpretation Based on Non-scientific Evidence and Inference?

The revision introduces the following (non)criteria for deciding that blasts or explosions caused skeletal trauma: "Blasts/explosive events often cause blunt (including concussive) and projectile trauma to the body. When the trauma pattern and circumstantial information support a blast event, the trauma mechanism should be classified as 'blast trauma'”. The undefined notion of "support" is too vague to give any guidance. Is "consistent with" considered "support"? Let's hope not -- patterns can be "consistent with" one hypotheses (it could occur when the hypothesis is true) but much more probable under the opposite hypothesis.

And then there is the green light this recommendation gives to presenting a conclusion based on nonscientific "circumstantial evidence" as if it were based on expertise involving the skeletal evidence. Knowing that a blast occurred can drive the conclusion that the damage to the skeleton is "blast trauma." Should there also be a report on the skeletal evidence from an analyst blinded to the other information uncovered in the investigation?

Is Everything Acceptable?

ASB 147-21 states that "Skeletal trauma shall be examined. Acceptable methods to examine trauma include gross, microscopic, radiographic, and other analytical methods." This formulation deems every conceivable analytical method as "acceptable" no matter how poorly conceived it may be. Labeling everything as "acceptable" is troublesome in a standard that does not include criteria and procedures for performing the analysis and that does not lead the reader to any evidence of the reliability and validity of the undefined "analytical procedures."

Of course, forensic anthropologists know that some procedures do not work well, and only an outlier would use them. The drafters of ASB 147-21 undoubtedly appreciate the need for suitable methods (and hence prohibit certain conclusions that cannot be drawn with any existing method). Well motivated and informed forensic anthropologists will not be led astray if they consult the standard. But outliers do appear in court. Remember Louise Robbins. Unless the dubious method yields one of the explicitly prohibited statements in this standard, the outlier witnesses could maintain that they have proceeded exactly as the standard requires. Standards with this potential for abuse should be reformed.They should strive to standardize the methods they govern, and they should state what is known about the accuracy and reliability of these methods.

Thursday, December 22, 2016

Realistically Testing Forensic Laboratory Performance in Houston

The Houston Forensic Science Center, announced on November 17, 2016, that
HFSC Begins Blind Testing in DNA, Latent Prints, National First
This innovation -- said to be unique among forensic laboratories and to exceed the demands of accreditation -- does not refer to blind testing of samples from crime scenes. It is generally recognized that analysts should be blinded to information that they do not need to reach conclusions about the similarities and differences in crime-scene samples and samples from suspects or other persons of interest. One would hope that many laboratories already employ this strategy for managing unwanted sources of possible cognitive bias.

Perhaps confusingly, the Houston lab's announcement refers to "'blindly' test[ing] its analysts and systems, assisting with the elimination of bias while also helping to catch issues that might exist in the processes." More clearly stated, "[u]nder HFSC’s blind testing program analysts in five sections do not know whether they are performing real casework or simply taking a test. The test materials are introduced into the workflow and arrive at the laboratory in the same manner as all other evidence and casework."

A month earlier, the National  Commission on Forensic Science unanimously recommended, as a research strategy, "introducing known-source samples into the routine flow of casework in a blinded manner, so that examiners do not know their performance is being studied." Of course, whether the purpose is research or instead what the Houston lab calls a "blind quality control program," the Commission noted that "highly challenging samples will be particularly valuable for helping examiners improve their skills." It is often said that existing proficiency testing programs not only fail to blind examiners to the fact that they are being tested, but also are only designed to test minimum levels of performance.

The Commission bent over backward to imply that the outcomes of the studies it proposed would not necessarily be admissible in litigation. It wrote that
To avoid unfairly impugning examiners and laboratories who participate in research on laboratory performance, judges should consider carefully whether to admit evidence regarding the occurrence or rate of error in research studies. If such evidence is admitted, it should only be under narrow circumstances and with careful explanation of the limitations of such data for establishing the probability of error in a given case.
The Commission's concern was that applying statistics from work with unusually difficult cases to more typical casework might overstate the probability of error in the less difficult cases. At the same time, its statement of views included a footnote implying that the defense should have access to the outcomes of performance tests:
[T]he results of performance testing may fall within the government’s disclosure obligations under Brady v Maryland, 373 U.S. 83 (1963). But the right of defendants to examine such evidence does not entail a right to present it in the courtroom in a misleading manner. The Commission is urging that courts give careful consideration to when and how the results of performance testing are admitted in evidence, not that courts deny defendants access to evidence that they have a constitutional right to review.
Using traditional proficiency test results and the newer performance tests in which examiners are blinded to the fact that they are being tested in a given case (which is a better way to test proficiency) to impeach a laboratory's reported results raises interesting questions of relevance under Federal Rules of Evidence 403 and 404. See, e.g., Edward J. Imwinkelried & David H. Kaye, DNA Typing: Emerging or Neglected Issues, 76 Wash. L. Rev. 413 (2001).

Friday, December 11, 2015

More on Task Relevance in Forensic Tests

Yesterday, I suggested that the National Commission on Forensic Science's views on task relevance are a significant step forward, and I elaborated on the use of conditional independence in determining which information is task relevant. The NCFS position is simple -- the examiner "should rely solely on task-relevant information when performing forensic analyses."

As the Commission explained, excluding task-irrelevant information avoids subtle but possible biases. 1/ However, what if the potentially biasing information could improve the accuracy of the analyst's conclusions? Statisticians often use biased estimators because they have greater precision -- they tend to give estimates that are closer to the true value with limited data -- even though these estimates tend to lie consistently on one side of that value. Moreover, even if the bias from the task-irrelevant information would increase the risk of an incorrect conclusion, what if it would be very costly to keep it out of the examination process? One might argue that the NCFS view is too stringent.

This challenge to the simple rule of no reliance is not persuasive. First, it is rather theoretical. Reasonably cheap methods to blind analysts to biasing, task-irrelevant information are generally available. The NCFS document explains how they can work.

Second, the conclusions that are likely to be more accurate are not those that the analyst should be drawing. At least with identification evidence in the courtroom, the expert should explain the strength of the scientific evidence, leaving the conclusion as to the identity of the true source to the judge or jury to decide based on all the evidence in the case. The NCFS views document adopts this philosophy most clearly in the last sentence of the appendix, which reads "[a]ny inferences analysts might draw from the task-irrelevant information involve matters beyond their scientific expertise that are more appropriately considered by others in the justice system, such as police, prosecutors, and jurors."

But it is not just a matter of relative expertise that should limit the analyst to task-relevant information. Forensic scientists are supposed to be conveying scientific information, and if the putative scientific judgment comes from a mixture of scientific and other information, the judge or jury cannot properly evaluate its weight without knowing what is the scientific part and what is some other part. 

Information contamination also makes it difficult to discern the validity of scientific tests. Consider hair-morphology evidence. I have presented the Houck-Budowle study of the correspondence between microscopic hair examinations and mitochondrial DNA tests as evidence that the former has some modest probative value (as measured by the likelihood ratio for positive associations). 2/ But inasmuch as the examiners were not blinded to task-irrelevant information, it is hard to tell from this one study how much of the probative value comes from the features of the hair and how much comes from other information that the hair examiners might have considered.

Studies of polygraphic lie detection offer another example. The technique sounds scientific, and the graphs of physiological responses look technical. But if the examiners' conclusions used in a validation study are influenced by impressions of the subject, the study does not reveal the diagnostic value of just the information in the tracings  -- the impact of that information and the subjective impressions are confounded. (This problem can be avoided by computerized scoring of the data.)

As the NCFS appendix emphasizes, the task-irrelevant information "does not help the analyst draw conclusions from the physical evidence that has been designated for examination through correct application of an accepted analytic method." At the risk of oversimplifying a complex subject, the message is that forensic scientists should stick to the scientific information.

Even this precept is not a complete response to concerns about bias. What if the task-relevant information also poses a serious risk of bias? If the contribution to the scientific analysis is minor and the risk of distortion is great, should not the examiner be blinded to this concededly task-relevant information? NCFS expressed no view on this situation. Perhaps it never arises, but if it does, standard-setting organizations should deal with it.

Notes
  1. The NCFS observes that "there are risks entailed in exposing examiners unnecessarily to task-irrelevant information." But if the information is truly task-irrelevant, why would it be necessary? And if such information exists, would not the same risk of biasing the analysis be present?
  2. David H. Kaye, Ultracrepidarianism in Forensic Science: The Hair Evidence Debacle, 72 Wash. & Lee L. Rev. Online 227 (2015); Disentangling Two Issues in the Hair Evidence Debacle, Forensic Sci., Stat. & L., Aug. 22, 2015, http://for-sci-law.blogspot.com/2015/08/disentangling-two-issues-in-hair.html.

Thursday, December 10, 2015

Blinding Forensic Analysts to Task-irrelevant Information: A National Commission (NCFS) Speaks Out

This week, the National Commission on Forensic Science (NCFS) approved a “views document” entitled Ensuring That Forensic Analysis Is Based Upon Task-Relevant Information. 1/ If these views are translated into practice, it will be a major step forward in making sure forensic science findings are based on scientific data and not on extraneous information. The document is thus cause for celebration.

Here, I describe how the document defines task-relevance. I identify an arguable inconsistency in the Commission’s terminology and elaborate on the use of what is known in probability theory as conditional independence.

The NCFS’ views are these:
  1. FSSPs [Forensic Science Service Providers] should rely solely on task-relevant information when performing forensic analyses.
  2. The standards and guidelines for forensic practice being developed by the Organization of Scientific Area Committees (OSAC) should specify what types of information are task-relevant and task-irrelevant for common forensic tasks.
  3. Forensic laboratories should take appropriate steps to avoid exposing analysts to task-irrelevant information through the use of context management procedures detailed in written policies and protocols.
The analysis and explication that follows this enumeration tries to define task-relevance both in words and in symbols involving conditional probabilities. The NCFS definition is in two parts:
(1) [I]nformation is task-relevant for analytic tasks if it is necessary for drawing conclusions: (i.) about the propositions in question, (ii.) from the physical evidence that has been designated for examination, (iii.) through the correct application of an accepted analytic method by a competent analyst.
(2) Information is task-irrelevant if it is not necessary for drawing conclusions about the propositions in question, if it assists only in drawing conclusions from something other than the physical evidence designated for examination, or if it assists only in drawing conclusions by some means other than an appropriate analytic method.
Taken literally, this formulation seems to dismiss as task-irrelevant information that could help the analyst assess the strength of the evidence yet is not necessary for drawing conclusions about the propositions.  For example, suppose the proposition P in question is whether a trace sample that has both clear and ambiguous features came from a suspect. Viewing a tape of someone who looks like (and thus might be) the suspect leaving the mark is task-irrelevant under (2). No analyst needs to view the tape to compare the questioned mark to a known exemplar. The analyst can reach a conclusion of some sort without the video.

Nevertheless, viewing the tape could help the analyst doing a side-by-side comparison resolve the ambiguities in the features in the mark and thereby “assess the strength of the inferential connection between the physical evidence being examined and the propositions the analyst is evaluating.” For example, if the mark is a distorted fingerprint, observing how it was deposited might help the analyst. It seems as if the tape should be declared task-relevant, but (1) requires that it be necessary to the analysis. Strictly speaking, it is not.

Indeed, a few sentences later, the document states that information “is task-relevant if it helps the analyst assess the strength of the inferential connection between the physical evidence being examined and the propositions the analyst is evaluating.” Not everything that is helpful is necessary.

That the Commission did not really mean to require necessity also can be gleaned from the “more formal definition of task-relevance and task-irrelevance ... in the technical appendix.” For “two mutually exclusive propositions P and NP that a forensic science service provider (FSSP) is asked to evaluate,” and for E defined as “the features or characteristics of the physical evidence that has been designated for examination,”
(1) information is task-relevant if it has the potential to assist the examiner in evaluating either the conditional probability of E under P—which can be written p(E|P)—or the conditional probability of E under NP—which can be written p(E|NP);
(2) information is task-irrelevant if it has no bearing on the conditional probabilities p(E|P) or p(E|NP).
Again, necessity is not crucial: The phrase “has the potential to assist” has been substituted for “is necessary,” and “has no bearing” has replaced “is not necessary.”

Technical definitions (1) and (2) also depart from (or refine) the main definitions (1) and (2) in that the only “conclusions” that can be considered in judging task-relevance are conditional probabilities for “features” given certain propositions. These conditional probabilities often are called “likelihoods” to distinguish them from the posterior probabilities of the propositions given the features. Traditionally, analysts testified about posterior probabilities expressed qualitatively or categorically. For example, the statement P that “Jane Doe’s thumb is the source of the latent print” is a categorical conclusion meaning that the posterior probability Pr(P|E) is close to 1.

Using likelihoods could be valuable in clarifying task-relevance, but the concepts of “bearing” and “potential to assist” remain undefined. It would seem that the NCFS intends to equate task-irrelevance with conditional independence. Let I denote the information that might be task-irrelevant. E and I are conditionally independent given some proposition R if and only if (iff) Pr(E&I|R) = P(E|R) P(I|R). An equivalent definition looks to whether Pr(E|I&R) = Pr(E|R). The idea is that once R is known, knowing I brings no additional information about E.

If we take conditional independence to be the NCFS’ technical definition of task-irrelevance, and we use R to stand for either P or NP, then we can rewrite the NCFS definitions as
(3) I is task-relevant iff Pr(E|I&R) ≠ Pr(E|R);
(4) I is task-irrelevant iff Pr(E|I&R) = Pr(E|R).
This more precise definition is easier to write than to apply. The “physical evidence” itself — bits of soil, specimens of handwriting, latent and rolled fingerprints, and so on — is not “E” in the probability function. Instead, E is “the features or characteristics of the physical evidence.” But are these the actual features or the declared features?

I think the formal definition works better when E refers to the true features (although the judge or jury only knows what the expert thinks they are). First, let’s look at an easy case. Suppose that E refers to the DNA alleles present at each locus in the suspect’s DNA (A0) and the profile in the crime-scene DNA (A1). Thus, E = A0 & A1. P means that the suspect is the source of both DNA samples; let Q mean that someone else is. I is a credible report that the suspect was near the crime scene just after the crime occurred. Finally, suppose that A0 and A1 are the same — the DNA in both samples have the same true features.

If P is true, then the samples must have the same features, so Pr(E|P) = Pr(E|I&P) =1. If Q is true, then whether the samples have the same features also does not depend on I — if someone else left the DNA, the suspect’s propinquity does not affect the alleles that the true contributor possesses and left at the crime-scene. Consequently, under (4), I is task-irrelevant, just as it should be.

Now, let’s make it more complicated. The laboratory is asked to assess whether a suspect’s “touch” DNA is present on a gun used in a killing. Several small peaks in the electropherogram are at the positions one would expect if this were the case, but they are at the limit of detectability. Some analysts would treat them as real (true peaks), but others would see them as spurious. The question is whether the analyst should be able to know the profile reported for the suspect — let’s call it r[A0] — before ascertaining the profile A1 in the crime-scene sample. Is I = r[A0] task-relevant to the determination of A1?

Some analysts might argue that I is task-relevant because knowing what is in the suspect’s DNA helps them understand what really is in the crime-scene DNA. They could say that the fact that the small peaks in the crime-science sample are located at just the same places as their larger homologs in the suspect’s sample helps them resolve the ambiguity arising from the small peak heights. Of course, if they are thoughtful, they also will recognize that the related information I could bias them, and they might well agree that they should not be exposed to it because it does not contribute enough to the accuracy of their determinations. But are they wrong in their claim that I is task-relevant (applying the NCFS definition)?

The views document does not give an explicit answer. It concludes with the observation that 
[Task-irrelevant information] might help the analyst draw conclusions about the propositions, but it does not help the analyst draw conclusions from the physical evidence that has been designated for examination through correct application of an accepted analytic method. Any inferences analysts might draw from the task-irrelevant information involve matters beyond their scientific expertise that are more appropriately considered by others in the justice system, such as police, prosecutors, and jurors.
This relative-expertise criterion, however, does not quite define task-irrelevance. The inference that a small peak in an electropherogram is the result of chemiluminescence from alleles as opposed to an artifact or background noise may be difficult to make correctly, but it is not clear that it is lies more squarely within the expertise of police, prosecutors, and jurors than of DNA analysts. 2/

The formal definition can help us out here. If P is true, then regardless of what the peaks look like and irrespective of the suspect’s reported profile r[A0], the true profile of the crime-scene sample is A1 = A0. Thus, Pr(E=A1|P) = Pr(A1|P&I) = Pr(A1|P&r[A0]) = 1. Likewise, if someone else’s DNA is on the gun instead of the suspect’s, then the probability that the profiles match also is unrelated to a report of what is in the suspect’s DNA sample. Once again, the conditional-independence definition of task-irrelevance seems to work.  Sometimes probability notation is purely window dressing, but the approach begun in the technical appendix might do some useful work in spotting task-irrelevant information. 3/

Notes
  1. The document should appear on the Commission’s web page http://www.justice.gov/ncfs/work-products-adopted-commission in the near future. The principal drafter of the document was Bill Thompson, who is the chair of the Human Factors Resource Committee of the NIST Organization of Scientific Area Committees (OSACs) that is developing standards for forensic science.
  2. One can argue that looking at the suspect's profile or peaks before resolving ambiguities in the crime-scene profile is not a "correct application of an accepted analytic method." However, given that the task is to ascertain and compare the two DNA samples, it seems odd to call this information about the profiles "irrelevant" as opposed to improper or not acceptable. And, if there were no standard in place rejecting this practice (as was true for a period of time), this criterion would not render the information task-irrelevant.
  3. This is so even though the likelihoods involved are not necessarily the ones that determine the probative value of the forensic analysis with respect to the two competing hypotheses P and Q. Those likelihoods are Pr(E*|P) and Pr(E*|Q). The asterisk is attached because we do not know the true features E in the samples. We have data E* on them (fallible measurements or observations of them). The probative value of E* (or, if you like, of the analysis that generates these data) is the likelihood ratio Pr(E*|P) / Pr(E*|Q). For example, even though we can speak of the likelihood ratio for the true DNA profiles (A1 & A0) for present purposes, the court’s evidence is the reported profiles: E* = r[A1 & A0].