Saturday, January 15, 2022

Bones of Contention: A Standard for Analyzing Skeletal Trauma in Forensic Anthropology

The Academy Standards Board (ASB) of the American Association of Forensic Sciences (AAFS) posted the second proposed draft of a "Standard for Analyzing Skeletal Trauma in Forensic Anthropology" for public comment. The standard does not go far toward standardizing procedures or showing that the procedures to which it applies have been scientifically tested. Of course, it could well be that ample, well designed studies have demonstrated that forensic anthropologists can consistently and accurately classify skeletal defects in human remains according to the categories the standard mentions. But the standard contains no bibliography and no citations to show that this is the case. 

It contains some negative injunctions and a few positive suggestions about reporting -- for example:

  • Forensic anthropologists shall not determine cause or manner of death.
  • Practitioners shall not estimate the temperature or duration of heat exposure based on thermal defects to bone.
  • Practitioners may report the minimum number of traumatic events (e.g., blunt impacts, projectile entry defects, or sharp defects) observed skeletally, but shall not report a definitive maximum number of impacts, as skeletal trauma evidence may not reflect all impacts to the body.
  • When a suspect tool is submitted for analysis, similarities between the tool and defect may be reported; conclusions shall be reported in terms of an exclusion or failure to exclude.

As such, ASB 147-21 is not without any redeeming legal value. Nevertheless, it does not articulate any analytical process by which the classifications it calls for should be made (cf. "vacuous standards"); it requires no reporting of the uncertainty in this process; it does not contemplate the possibility of evidence-based rather than conclusion-based statements of the implications of the data; and it refers to an all-inclusive list of methods as "acceptable." If I may elaborate:

File:Human skeleton remains.jpg - Wikimedia Commons

Is "Interpretation" Limited to an Opinion on the Inference (Conclusion) from the Data?

The revision defines "trauma interpretation" as "Opinion regarding the mechanism of, timing, direction of impact(s) or minimum number of impacts associated with skeletal defect(s) based on quantitative and/or qualitative observations." The phrase "based on ... observations" indicates that the opinion expresses a belief in the truth, falsity, or probability of an inference being drawn from the data. Interpretation should include the possibility of describing the strength of the evidence in favor of the inference rather than opining on the truth, falsity, or probability of the conclusion itself. In addition, if the opinion-statement is an assertion that the hypothesis about what happened is true or false (either categorically or to some probability), it is not just based on the data but on a prior probability for the hypothesis as well.

Despite these definitions, the standard sanctions "interpretation" in the form of rudimentary statements about the extent to which the data prove the hypothesis in question. Section 6 notes that "Trauma interpretation shall be clearly identified in the report using terms such as ‘indicative of’ and ‘consistent with’ or by using a subheading titled ‘Interpretation.’"These phrases have their problems, but they are one manner of referring to the probability of the evidence given the truth of certain probabilities rather than vice versa.

Is Interpretation Based on Non-scientific Evidence and Inference?

The revision introduces the following (non)criteria for deciding that blasts or explosions caused skeletal trauma: "Blasts/explosive events often cause blunt (including concussive) and projectile trauma to the body. When the trauma pattern and circumstantial information support a blast event, the trauma mechanism should be classified as 'blast trauma'”. The undefined notion of "support" is too vague to give any guidance. Is "consistent with" considered "support"? Let's hope not -- patterns can be "consistent with" one hypotheses (it could occur when the hypothesis is true) but much more probable under the opposite hypothesis.

And then there is the green light this recommendation gives to presenting a conclusion based on nonscientific "circumstantial evidence" as if it were based on expertise involving the skeletal evidence. Knowing that a blast occurred can drive the conclusion that the damage to the skeleton is "blast trauma." Should there also be a report on the skeletal evidence from an analyst blinded to the other information uncovered in the investigation?

Is Everything Acceptable?

ASB 147-21 states that "Skeletal trauma shall be examined. Acceptable methods to examine trauma include gross, microscopic, radiographic, and other analytical methods." This formulation deems every conceivable analytical method as "acceptable" no matter how poorly conceived it may be. Labeling everything as "acceptable" is troublesome in a standard that does not include criteria and procedures for performing the analysis and that does not lead the reader to any evidence of the reliability and validity of the undefined "analytical procedures."

Of course, forensic anthropologists know that some procedures do not work well, and only an outlier would use them. The drafters of ASB 147-21 undoubtedly appreciate the need for suitable methods (and hence prohibit certain conclusions that cannot be drawn with any existing method). Well motivated and informed forensic anthropologists will not be led astray if they consult the standard. But outliers do appear in court. Remember Louise Robbins. Unless the dubious method yields one of the explicitly prohibited statements in this standard, the outlier witnesses could maintain that they have proceeded exactly as the standard requires. Standards with this potential for abuse should be reformed.They should strive to standardize the methods they govern, and they should state what is known about the accuracy and reliability of these methods.

Monday, January 3, 2022

Fitting "Physical Fit" into the Courtroom

The logic of piecing together fragments of broken glass, torn tape, cut paper, and the like seems simple enough. \1/ If the pieces fit in all their details at the edges, and if all surface marks or impressions that would cross an edge also align nicely, one has circumstantial evidence that they were once part of the same object.

The strength of this evidence for a single source depends on the extent and detail of the concordance between the recovered pieces. A physical fit between two halves of a broken plank of wood is powerful evidence for the hypothesis that the two pieces resulted from breaking this one plank. But if the pieces are weathered and the splintered edges dulled, the physical fit will be less precise and less supportive of the claim that they came from the same original plank.

At the other extreme, if two pieces are plainly discordant, they might have come from different places on the same object, with the intermediate pieces being missing. Or they might have come from different objects entirely. Consider tearing off five pieces of duct tape from the same roll of tape and comparing the edges of the first and the last segments. The detailed structure of the edges should not be complementary. Likewise, tearing segments of tape from five different rolls should result in a mismatch between the first and the fifth segment.

Criminalists or materials experts can be extremely helpful in examining the recovered pieces of objects to determine the degree of physical fit -- that is, in elucidating how well the edges fit together and the extent to which a mark on the surface of one piece lines up with a mark on the other when the pieces are aligned. But how they should describe their findings seems to be muddled in forensic-science standards. This posting describes the current vocabulary and argues that it is articificial and a departure from the ordinary meaning of the term "fit." It then outlines better alternatives to reporting the results of an investigation into physical fit.

I. The Standard Approach

Let’s look at a couple of ASTM standards. E2225-19a (Standard Guide for Forensic Examination of Fabrics and Cordage) instructs that “[i]f a physical match is found, it should be reported in a manner that will demonstrate that the two or more pieces of material were at one time a continuous piece of fabric or cordage” (§ 7.2.2). This standard treats the “physical match” as an observable property of the specimens (concordant edges and surface marks) that is conclusive of the hypothesis of a single source (the inference from the data).

ASTM E3260−21 (Standard Guide for Forensic Examination and Comparison of Pressure Sensitive Tapes), on the other hand, characterizes “physical fit” not as a property of the materials, but as a “type of examination that can be performed” (§ 10.5.1). This “conclusive type of examination ... is a physical end match.” Id. It “involves the comparison of edges, fabric (if present), surface striae, and other surface irregularities between samples in which corresponding features provide distinct characteristics that indicate the samples were once joined at the respective separated edges.” Of course, "distinct characteristics that indicate the samples were once joined at the respective separated edges” are not necessarily "conclusive," making this definition of "physical fit" as a "type of examination" puzzling. The intent, it seems, is to define a physical fit examination (rather than a physical fit) as one that is capable of conclusively proving that the pieces were once joined together.

A Proposed New Standard Guide for the Collection, Analysis and Comparison of Forensic Glass Samples, ASTM WK72932, released for public comment late last year states that “broken objects can be reassembled to their original configuration ... called a ‘physical fit’ (§ 11.1). But a physical fit is the original configuration of a broken object only if the pieces come from that original object, and this origin story is not true just because a standard defines "physical fit" that way. The evidence from the examination may be that the separate pieces fit together extremely well. If so, the conclusion is that they were once together within or as a unitary object. This conclusion may well be true, but one cannot decide, by the fiat of a definition, that the pieces that are observed to fit together well have been realigned as they once were. Yet, a later section similarly asserts that “[a] glass physical fit is a determination that two or more pieces of glass were once part of the same broken glass object” (§ 11.2.8). This effort to define "physical fit" as inherently conclusive prompted eleven lawyers (including me) \2/ to caution ASTM that “[t]he hypothesis or conclusion that fragments come from the same object is not a physical fit. It is an inference drawn from the observations that produce the designation of a physical fit.”

Still more recently, an OSAC subcommittee released a Standard Guide for Forensic Physical Fit Examination (OSAC 2022-S-0015) for public comment before it is delivered to ASTM for consideration there. This proposed standard goes off in another direction. It equates a “physical fit” with the examiner’s state of mind about a hypothetical ensemble of experiments:

13.1 Physical Fit
13.1.1 The items that have been broken, torn, separated, or cut exhibit physical features that realign in a manner that is not expected to be replicated.
13.1.1.1 Physical Fit is the highest degree of association between items. It is the opinion that the observations provide the strongest support for the proposition that the items originated from the same source as opposed to the proposition they originated from different sources.

13.2 No Physical Fit
13.2.1 The items correspond in observed class characteristics, but exhibit physical features that do not realign, or they realign in a manner that could be replicated.
13.2.2 Alternatively, the items can exhibit physical features that partially realign, display simultaneous similarities and differences, show areas of discrepancy (e.g., warped areas, burned areas, missing pieces), or have insufficient individual characteristics that hinder the ability to determine the presence or absence of a physical fit.

Statisticians will notice the shift from (1) the incompletely expressed frequentist idea of an infinite sequence of trials in which different objects A and B are broken and the pieces from A never align with those from B to (2) the likelihoodist conception of support for the same-source hypothesis. But that implicit change in the theory of inference is hardly a cardinal sin in this context. If the probability of a fit at least as good as the one observed is practically zero for different sources, and if the probability of such a fit for the same source is much higher, then the support (the log-likelihood ratio) is very high.

Nevertheless, defining physical fit as a categorical opinion rather than a more variable degree of congruency that generates the opinion — and dumping everything short of a perceived fit into the category of ”no physical fit” — deviates from the common understanding that physical fit comes in degrees. There can be a remarkably great fit, a pretty good fit, and so on, down to a blatant misfit. The question the examiner must answer, at least intuitively, before the fit/no-fit classification can be made is just how well the pieces fit together. Fit is not a uniform degree of association that springs into existence exactly when a particular examiner is convinced that no other source could account for the complexity and extent of the fit. There is no such thing as “the strongest support.” One can always conceive of a situation with still stronger support (because a fracture or other separation of the pieces could generate an even richer set of irregularities in the edges).

The current approach of defining a physical fit as a single source for the pieces and calling everything else “no fit” does not create a vocabulary that judges or jurors will easily understand. A vocabulary in which physical congruency (fit) lies on a continuum — and that then addresses the inference that should be drawn from the observations — is more transparent.The definitions in the standards collapse the two steps of data acquisition and inference into one.

II. Inference: From Data to Conclusions

So how should examiners answer the question of how well the pieces fit together? An examination for fit yields multidimensional, spatial data. An examiner could present photographs of the aligned edges and surfaces and highlight the concordant and discordant features. Although the highlighting involves some interpretative thinking, I have called a courtroom presentation that stops at this point "features-only testimony." \3/ It is appropriate when examiners have no special expertise at interpreting how strongly their results support the same-source hypothesis. If they are no better than lay judges and jurors at discerning how improbable the features are in the hypothetical cases of repeatedly breaking the same object, it could be argued that these witnesses should not try to interpret the results any further. Such interpretation would not actually assist the trier of fact, as required by Federal Rule of Evidence 702.

For example, a few days ago, a forensic scientist told me of a case in which a criminalist was able to reassemble pieces of glass recovered at the site of a hit-and-run accident so that they fit neatly into the metal holder of a side rear mirror on the suspect’s car that was missing its glass. That’s good detective work, but did the criminalist have any special insights to offer into the obvious implications of this solution to the jigsaw puzzle? (The work was not presented in court because the crime laboratory’s management was concerned that there was no written protocol for pasting mirror fragments back in place. As the scientist observed, that's silly. The evidence practically speaks for itself, and its message is the same with or without a written protocol.)

Nevertheless, let’s assume that examiners do have specialized skill at interpreting the findings about the alignment of the features. The ASTM and OSAC-proposed standards ignore the possibility of a qualitative expression of relative support — for example, “It is far more likely to get the detailed alignment of the features I just showed you if the pieces were broken parts of the same objects than if they were from different objects.” Or, similarly, “The detailed alignment gives very strong support to the idea that the pieces broke off of the same object as opposed to two different objects.”

As Part I showed, the standards advocate a fit/no-fit classification in which “fit” is either a statement about the probability of the same-source hypothesis (that the pieces had to have come from the same object) or a statement of belief in the hypothesis (“my opinion is that they were together in the same object — that’s what makes it a physical fit). No-fit does not have a comparably sharp meaning. It could mean anything from no realistic possibility that the pieces were once contiguous parts of the same object to “partial fit features [that] increase the significance of the finding” (OSAC 2022-S-0015 § 13.2.4).

A more straightforward and comprehensible approach would be to have a three-tiered reporting scale for the support the data give to the same-source hypothesis. What is now called a physical fit would be designated a highly probative physical fit (that is, a physical fit that strongly supports the same-source hypothesis). “Partial fit features” would be described as a limited fit (that gives some support to the same-source hypothesis). Finally, an obvious mismatch could be called a misfit (which strongly supports the conclusion that the pieces were never adjacently located on the same object).

This tripartite classification is an imperfect way to express an underlying likelihood ratio formed from subjective probabilities. Whether better results would be achieved if analysts were forced to articulate their probabilities, either quantitatively or in the qualitative way mentioned earlier, is an interesting question. But the three-tiered reporting scale is closer to the current practice and seems feasible. \4/ It offers a framework for a better standard on reporting the results of a physical fit examination. Or so it seems to me — those who disagree are encouraged to hit the comment button.

NOTE

  1. But see Forensic Science’s Latest Proof of Uniqueness, Dec. 22, 2013, http://for-sci-law.blogspot.com/2013/12/forensic-sciences-latest-proof-of.html.
  2. The other commenters were Alyse Bertenthal, Amanda Black, Jennifer Friedman, Julia Leighton, Kate Philpott, Emily Prokesch, Matt Redle, Andrea Roth, Maneka Sinha, and Pate Skene.
  3. David H. Kaye et al., The New Wigmore on Evidence" Expert Evidence (2d ed. 2011).
  4. When there is a mismatch, testimony about a physical match has little value. Other features than the alignment of edges and surface markings will need to be studied if the expert is to shed light on whether the pieces came from a single object. The current and proposed standards are clear on this point.

Saturday, December 25, 2021

The FBI's Misinformation Campaign on Firearms-toolmark Testimony

On Tuesday (21 December 2021), the Texas Forensic Science Commission issued a Statement Regarding 'Alternate Firearms Opinion Terminology'. It is a forceful correction to misinformation from the FBI Laboratory's Assistant General Counsel, Jim Agar II. \1/ The email that attracted the Commission's critical attention tells forensic analysts what they are supposed to say in opposition to motions to limit their testimony about firearms-toolmark comparisons. As previous postings show, there has been no shortage of defense motions seeking to forbid eliciting opinions that ammunition components associated with a crime came from a particular gun.

The FBI advice to firearms examiners is entitled "Dealing with Alternate Firearms Opinion Terminology" (hereinafter Dealing). It begins by dismissing the best efforts of federal and state judges to respond to weaknesses in traditional "This is the gun!" testimony as "wholesale attempts to rewrite the firearm expert's testimony by a layman with no experience in forensic science." \2/ The fact that eminent scientists and respected jurists have questioned source-attribution testimony in general and in this field in particular does not seem to matter. According to Dealing, the limitations are "not supported by either science or the law." Despite the government's annoyance with lay judges' rulings, however, courts have a duty to review the scientific and scholarly literature to decide whether strong claims of source attributions are sufficiently warranted. \3/

Dealing continues, more reasonably, with the strategic recommendation that "firearms examiners and prosecutors should address the terminology issue head-on during their direct examination at the admissibility hearing. Preempt this issue early. Don't wait for the judge or the defense counsel to bring it up." But the tactics for bringing it up are over the top. Dealing imagines the following colloquy:

Prosecutor: Can you testify truthfully that your opinion is that the cartridge cases and/or bullets in this case
   • "Could or may have been fired by this gun?"
   • ''Are consistent with having been fired by this gun?"
   • "Are more likely than not having been fired by this gun?"
   • "Cannot be excluded as having been fired by this gun?"
Examiner: No, I cannot testify truthfully to any of those statements or just the class characteristics alone.
Prosecutor: Why not?
Examiner: For three reasons: First, there are no empirical studies or science to backup any of those statements or terminology. Second, those statements are not endorsed nor approved by my laboratory, any nationally recognized forensic science organization, law enforcement, or the Department of Justice. Third, those statements are false as they do not reflect my true opinion of identification. Such statements would mislead the jury about my opinion in this case. It would also constitute a substantive and material change to my opinion from one of Identification to Inconclusive. This would constitute perjury on my part for I would not be telling the jury the whole truth.

The "three reasons" border on the absurd (if they do not cross the border). First, the empirical studies that prosecutors cite to support the ability of firearms experts to match ammunition components to specific guns also support the bulleted statements. This is because the alternatives are lesser included statements, so to speak. If a categorical source attribution is correct, then a weaker included statement such as "cannot be excluded" also is true. If "empirical studies or science" do not adequately support these weaker statements, then, a fortiori, they do not support the much stronger claims that Dealing advocates.

Second, that law enforcement organizations and crime laboratories do not approve of the policy of replacing traditional "This is the gun!" testimony with a less telling alternative proves nothing about whether the bulleted statements are true or false. It merely means that a laboratory is unwilling to change its standard operating procedure and that "law enforcement" opposes losing the opinions that prosecutor's love their experts to provide. No self-respecting expert can say that the desire of "law enforcement" and crime laboratories for the strongest possible testimony makes less compelling testimony "untruthful."

Finally, that any lawyer -- let alone one representing the FBI -- would ask a forensic examiner to tell a judge that it would be perjurious to testify in the bulleted ways is shocking. A federal perjury prosecution would be laughed out of court. Under federal law, statements that are known to be incomplete, or, worse, fully intended to distract or mislead, do not constitute perjury if they are literally true. The leading case is Bronson v. United States. \4/ There, defendant testified as follows:

Q. Do you have any bank accounts in Swiss banks, Mr. Bronston?
A. No, sir.
Q. Have you ever?
A. The company had an account there for about six months, in Zurich.
Q. Have you any nominees who have bank accounts in Swiss banks?
A. No, sir.
Q. Have you ever?
A. No, sir.

In reality, the witness had previously maintained and had made deposits to and withdrawals from a personal bank account in Geneva, Switzerland. Clearly, his answers were calculated to avoid revealing this fact. However, the Supreme Court unanimously reversed a conviction for perjury, concluding that the federal statute did not criminalize lying by omission and misdirection.

To be sure, some state statutes define the crime to encompass wilful omissions, but the core idea remains that perjury occurs when the witness intends to give the questioner false information or a false impression so as to obstruct the ascertainment of the truth. \5/ An expert witness who testifies sincerely to true statements such as "the defendant's gun cannot be excluded as the one that fired the recovered bullet" or "measurements of the bullet and the pistol showed them both to be 9 mm, so the bullet could have been fired from the gun," is not intending to lead anyone to a false conclusion. That the FBI would like firearms examiners to give more incriminating opinions does not make the lesser included testimony false or misleading. A prosecutor who truly is worried that "[t]estimony about class characteristics alone may falsely imply an examiner was unable to reach a conclusion of identification" can ask the court to instruct the jurors that the rules of evidence no longer allow an expert witness to testify that a bullet came from a particular gun and that they may not draw any inference from the absence of such inadmissible testimony. Instead, they are to use only the testimony that the expert gave in coming to a conclusion about which gun fired the recovered bullet.

After maintaining that "laymen" (courts) are asking toolmark examiners to commit perjury, Dealing gives another specious argument to persuade toolmark experts to stick to their guns (sorry about that) and refuse to "agree to testify to the terms of 'Could or may have fired,' or 'Consistent with,' 'More likely than not,' or 'Cannot be excluded.'" FBI counsel believes that examiners who testify this way when they feel that a traditional source attribution is justified "are ratifying these bogus statements and adopting this as their testimony, giving the judge a pass on the difficult decision to admit or exclude their testimony. They are also acquiescing to the judge's faulty terminology."

This is nonsense. The law has a spectrum of options ranging from excluding every bit of information a firearms expert might provide (which is unjustified given what is known about the performance of these experts) to unfettered admission of "This is the gun!" testimony (which is traditional). The only "fault" in the intermediate testimony is that it is not as strong as a prosecutor might want it to be. It is conservative in the sense of understating probative value (as FBI counsel understands the science), but testifying conservatively at trial when that is what a court requires does not "ratify" anything about the court's ruling. It simply presents a permissible opinion. DNA experts who testified to "ceiling" probabilities of random matches because that was the best the prosecution could get some courts to accept circa 1995 were not perceived as "ratifying these bogus statements." \6/

Dealing disagrees. FBI counsel insists that "acquiescing" in court rulings is "fatal" to an examiner's career as a witness:

This is fatal. Why? Once you testify to these bogus terms, you are wedded to them for life. At subsequent trials, defense counsel will pull out the verbatim transcript of the examiner's previous testimony where they used these court-induced terms. On cross examination, they will confront the examiner with their previous testimony and contrast their opinion of "Identification" with those in previous cases, then claim the expert is merely making this stuff up. The examiner no longer has any credibility in the jury's eyes.

This fear of cross-examination is fanciful. If the expert testifies at the admissibility stage (as Dealing contemplates) that "This is the gun!" testimony is scientifically justified, then that is what the expert is on record as stating. Later, more circumscribed testimony pursuant to court order is not an inconsistent statement useful for impeachment. Any competent expert witness will have no trouble explaining that in the earlier case, I reached the conclusion of "identification" (just read my case notes), and I used other terminology only because the prosecutor asking the question (or the judge) said I had to use the lesser included language because of a legal rule rather than a scientific principle.

In contrast, the witness who follows FBI counsel's advice will lose all credibility. The truth is that the lesser included testimony, while less powerful, is no less truthful than "This is the gun!" testimony. It is somewhat like choosing a wider confidence interval to increase the coverage probability; the statement becomes less precise, but it is more likely to be true. Talk of perjury and being asked to lie suggests either that (1) the witness does not understand a statement such as "the recovered bullet could have come from/is consistent with coming from/is not excluded as coming from/is more likely to have come from the firearm in question or that (2) the witness has chosen to lobby for the prosecution rather than to educate the judge impartially.

NOTES

  1. Mr. Agar is a decorated, retired Colonel with "31 years of successful experience leading complex legal organizations as a general counsel, attorney, leader, mentor and trainer of FBI legal offices and senior-level Army staffs" and "hands-on experience in advising senior FBI and Army leaders in all legal matters." His work as Assistant General Counsel for the "FBI Forensic Laboratory" began in October 2016. On Linkedin, from which these quotations are taken, he summarizes his current position as
    Legal advisor to the largest and best forensic laboratory in the world with a staff of over 700 scientists and a budget of $110 million. Responsible for training and qualifying the FBI’s forensic examiners to testify in any and all courts nationwide and internationally, consisting of over 120 examiners in 37 different disciplines. Coordinate all discovery for the Laboratory. Provide ethics advice to Laboratory personnel.
  2. Discussion of this line of cases can be found in David H. Kaye et al., Wigmore on Evidence: Expert Evidence (3d ed. 2021).
  3. The track record of the courts in translating this literature and the growing research on firearms-toolmark comparisons into appropriate constraints on proposed expert testimony is not perfect. Indeed, most of the judicial palliatives for perceived expert overclaiming (such as the supposed limitation of "a reasonable degree of ballistic certainty" and the alternatives listed in Dealing) are far from optimal. Id. (and other postings in this blog). But these failures hardly mean that, as "laymen," judges are disqualified from trying to improve the presentation of expert knowledge by excluding certain forms of testimony.
  4. Bronston v. United States, 409 U.S. 352 (1973).
  5. See Ira P. Robbins, Perjury by Omission, 97 Wash. U. L. Rev. 265 (2019).
  6. See, e.g., David H. Kaye, The Double Helix and the Law of Evidence (2010).

Friday, September 3, 2021

Does Qualitative Measurement Uncertainty Exist?

I have heard it said that forensic-science standards for interpreting the results of chemical or other tests need not discuss uncertainty in measurements of qualitative properties. For instance, ASTM International appropriately requires standards for test methods to include a section reporting on precision and bias as manifested in interlaboratory tests. Yet, it applies this requirement exclusively to quantitative measurements. Its 2021 style manual is unequivocal:

When a test method specifies that a test result is a nonnumerical report of success or failure or other categorization or classification based on criteria specified in the procedure, use a statement on precision and bias such as the following: “Precision and Bias—No information is presented about either the precision or bias of Test Method X0000 for measuring (insert here the name of the property) since the test result is nonquantitative" (ASTM 2020, § A21.5.4, pp. A3-A14).

Qualitative measurements are observation-statements such as the ink is blue, the friction ridge skin pattern includes loops, the bloodstain displays a cessation pattern, the blood group is type A, the glass fragments fit together perfectly, or the material contains cocaine. Likewise, the statements could be comparative: the recording of an unknown bell ringing sounds like it has a higher pitch than the ringing of a known bell; the hairs are microscopically indistinguishable; or the striations on the recovered bullet and the test bullet line up when viewed in the comparison microscope.

“Precision” is defined as “the closeness of agreement between test results obtained under prescribed conditions” (ibid. § A21.2.1, at A12). “A statement on precision allows potential users of the test method to assess in general terms its usefulness in proposed applications” and is mandatory (ibid. § A21.2, at A12). So how can it be that statements of precision and bias are not allowed for qualitative as opposed to quantitative findings? In both situations, the system that generates the findings could be noisy or skewed in its outcomes.

The only answer I have heard is that measurements cannot be qualitative because the word "measurement" is reserved for determining the magnitude of quantities such as length or mass. The values of these quantitative variables are basically isomorphic to the nonnegative real numbers. Counts, such as the number of alpha particles emitted in a given interval of time by radium atoms, also qualify as measurements because there is a quantitative, additive structure to them. The values of the variable are basically isomorphic to the natural numbers. Properties that only have names are described by nominal variables. Although numbers can assigned (1 for a match and 0 for a nonmatch, for example) these numbers are no more a measurement than a social security number is. In short, the argument is that because “measurements” do no not include qualitative judgments, classifications, decisions, identifications, or whatever one might call them, no statement of measurement uncertainty or error is possible, let alone required.

This argument is incredibly weak. To begin with, the definition of “measurement” is a highly contested concept. As one guide from NIST explains, a “much wider” conception of measurement than the one “contemplated in the current version of the International vocabulary of metrology (VIM)” has been developed in the metrology literature, and the measurand “may be ... qualitative (for example, the provenance of a glass fragment determined in a forensic investigation" (Possolo 2015). Broader conceptions of measurement have been the subject of many decades of writing in psychology and psychometrics (see, e.g., Humphry 2017; Mitchell 1990). Philosophers have been struggling to describe the scope and meaning of "measurement" at least since Aristotle (see, e.g., Tal 2015).

Second, even if one agrees with the definition in one NIST publication that “[m]easurement is [confined to] an experimental process that produces a value that can reasonably be attributed to a quantitative property of a phenomenon, body, or substance” (NIST 2019), some qualitative observations fit this definition. The color of a strip of litmus paper, for instance, can be understood as a value “that can reasonably be attributed to a quantitative property,” It is simply a crude measurement of pH.

Finally, the argument that there can be no measurement error for qualitative properties because those properties are not really “measured” is a semantic ploy that misses the point. The observations or estimates of nonquantitative properties as well as the individual measurements of quantitative properties are all subject to possible random and systematic error, and statements expressing the range of probable error for all measurements, observations, estimates, and classifications are essential. The need for these statements cannot be avoided for qualitative properties or judgments by the fiat of the VIM or some other dictionary. Even if “measurement” must be read in one particular, narrow, technical sense, “evaluation uncertainty” or “examination uncertainty” still must be reckoned with (Mari et al. 2020).

In sum, there is no excuse for ASTM and other organizations promulgating standards for forensic-science test methods to exempt any reported findings from required statements of uncertainty. Many statistics can be used to indicate how reliable (repeatable and reproducible) and valid (accurate) the test results may be (ibid.; Ellison & Gregory 1998; Pendrill & Petersson 2016). The qualitative-quantitative distinction affects the choice of the statistical method or expression but not the need to have one.

REFERENCES

  • ASTM Int’l, Form and Style for ASTM Standards (2020), https://www.astm.org/FormStyle_for_ASTM_STDS.html.
  • Stephen L. R. Ellison & Soumi Gregory, Perspective: Quantifying Uncertainty in Qualitative Analysis, Analyst 123, 1155-1161 (1998), https://doi.org/10.1039/A707970B
  • Stephen M. Humphry, Psychological Measurement: Theory, Paradoxes, and Prototypes, 27(3) Theory & Psychology 407–418 (2017)
  • L. Mari, C. Narduzzi, S. Trapmann, Foundations of Uncertainty in Evaluation of Nominal properties, 152 Measurement 107397 (2020), DOI:10.1016/j.measurement.2019.107397
  • Joel Mitchell, An Introduction to the Logic of Psychological Measurement (1990)
  • NIST, Statistical Engineering Division, Measurement Uncertainty, updated Nov. 15, 2019, https://www.nist.gov/itl/sed/topic-areas/measurement-uncertainty
  • Leslie Pendrill & Niclas Petersson, Metrology of human-based and other qualitative measurements, 27(9) Measurement Sci. Technol. 27 094003 (2016)
  • A. Possolo, Simple Guide for Evaluating and Expressing the Uncertainty of NIST
    Measurement Results (NIST Technical Note 1900), 2015, doi: 10.6028/NIST.TN.1900
  • Eran Tal, Measurement in Science, in Stanford Encyclopedia of Philosophy (Edward N. Zalta ed. 2015), https://plato.stanford.edu/archives/fall2017/entries/measurement-science/

APPENDIX: ADDITIONAL PUBLICATIONS ON "QUALITATIVE MEASUREMENT"

  1. Mary J. Allen & Wendy M. Yen, Introduction to Measurement Theory 2 (1979) ("In measurement, numbers are assigned systematically and can be of various forms. For example, labeling people with red hair "1" and people with brown hair "2" is a measurement. Since numbers are assigned to individuals in a systematic way and differences between scores represent differences in the property being measured (hair color).")
  2. Peter-Th. Wilrich, The determination of precision of qualitative measurement methods by interlaboratory experiments, Accreditation and quality assurance, 15: 439-444 (2010)
  3. Boris L. Milman, Identification of chemical compounds, Trends in Analytical Chemistry, 24:6, 2005 ("identification itself is considered as measurement on a qualitative scale")
  4. NIST Expert Working Group on Human Factors in Latent Print Analysis, Latent Print Examination and Human Factors: Improving the Practice Through a Systems Approach, Gaithersburg: National Institute of Standards and Technology, David H. Kaye ed., 2012 (defining "measurement" broadly, to encompass categorical variables, including the examiner's judgment about the source of a print).
  5. Lim, Yong Kwan, Kweon, Oh Joo, Lee, Mi-Kyung and Kim, Hye Ryoun. Assessing the measurement uncertainty of qualitative analysis in the clinical laboratory. Journal of Laboratory Medicine, vol. 44, no. 1, 2020, pp. 3-10. https://doi.org/10.1515/labmed-2019-0155 ("Measurement uncertainty is a parameter that is associated with the dispersion of measurements. Assessment of the measurement uncertainty is recommended in qualitative analyses in clinical laboratories; however, the measurement uncertainty of qualitative tests has been neglected despite the introduction of many adequate methods.")
  6. Donald Richards, Simultaneous Quantitative and Qualitative Measurements in Drug-Metabolism Investigations, Pharmaceutical Technology 2013
  7. Kadri Orro, Olga Smirnova, Jelena Arshavskaja, Kristiina Salk, Anne Meikas, Susan Pihelgas, Reet Rumvolt, Külli Kingo, Aram Kazarjan, Toomas Neuman & Pieter Spee, Development of TAP, a non-invasive test for -qualitative and quantitative measurements of biomarkers from the skin surface, Biomarker Research 2: 20 (2014)
  8. J M Conly & K Stein, Quantitative and qualitative measurements of K vitamins in human intestinal contents, Am J Gastroenterol. 1992 Mar;87(3):311-316
  9. Wenjia Meng, Qian Zheng, Gang Pan, Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network, IEEE Transactions on Neural Networks and Learning Systems 2020
  10. Rudolf M. Verdaasdonk, Jovanie Razafindrakoto, Philip Green, Real time large scale air flow imaging for qualitative measurements in view of infection control in the OR (Conference Presentation) Proceedings Volume 10870, Design and Quality for Biomedical Technologies XII; 1087002 (2019) https://doi.org/10.1117/12.2511185
  11. Rashis, Bernard, Witte, William G. & Hopko, Russell N., Qualitative Measurements of the Effective Heats of Ablation of Several Materials in Supersonic Air Jets at Stagnation Temperatures Up to 11,000 Degrees F, National Advisory Committee for Aeronautics, July 7, 1958
  12. Lawrence F Cunningham and Clifford E Young, Quantitative and Qualitative Approaches, Journal of Public Transportation 1(4) (1997) ("The study also contrasts the results of quantitative and qualitative measurements and methodologies for assessing transportation service quality")
  13. JM Conly, K Stein, Quantitative and qualitative measurements of K vitamins in human intestinal contents, American Journal of Gastroenterology, 1992
  14. P Sinha, Workshop on Biologically Motivated Computer Vision, 2002 - Springer ("Our emphasis on the use of qualitative measurements renders the representations stable in the presence of sensor noise and significant changes in object appearance. We develop our ideas in the context of the task of face-detection under varying illumination")
  15. D Michalski, S Liebig, E Thomae & A Hinz, Pain in Patients with Multiple Sclerosis: a Complex Assessment Including Quantitative and Qualitative Measurements, 40 J. Pain 219–225 (2011), https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3160835/
  16. Cécilia Merlen, Marie Verriele, Sabine Crunaire,Vincent Ricard, Pascal Kaluzny, Nadine Locoge, Quantitative or Only Qualitative Measurements of Sulfur Compounds in Ambient Air at Ppb Level? Uncertainties Assessment for Active Sampling with Tenax TA®, 132 Microchemical J. 143-153 (2017)
  17. Tomomichi Suzuki, Jun Ichi Takeshita, Mayu Ogawa, Xiao-Nan Lu, Yoshikazu Ojima, Analysis of Measurement Precision Experiment with Categorical Variables, 13th International Workshop on Intelligent Statistical Quality Control 2019, Hong Kong ("Evaluating performance of a measurement method is essential in metrology. Concepts of repeatability and reproducibility are introduced in ISO5725-1 (1994) including how to run and analyse experiments (usually collaborative studies) to obtain these precision measures. ISO5725-2 (1994) describe precision evaluation in quantitative measurements but not in qualitative measurements. Some methods have been proposed for qualitative measurements cases such as Wilrich (2010), de Mast & van Wieringen (2010), Bashkansky, Gadrich & Kuselman (2012). Item response theory (Muraki, 1992) is another methodology that can be used to analyse qualitative data.").

Monday, June 14, 2021

Tibbs. Shipp, and Harris on "Meaningul" Peer Review of Studies on Firearms-toolmark Matching

The Supreme Court's celebrated (but ambiguous) opinion in Daubert v. Merell Dow Pharmaceuticals, \1/ was a direct response to a seemingly simple rule--results that are not published in the peer-reviewed scientific literature are inadmissible to prove that a scientific theory or method is generally accepted in the scientific community. The Court unanimously rejected this strict rule--and more broadly, the very requirement of general acceptance--in favor of a multifaceted examination guided by four or five criteria that have come to be known as "the Daubert factors."

But "peer review and publication" lives on--not as a formal requirement, but as one of these factors. Thus, courts routinely ask whether the peer-reviewed scientific literature supports the reasoning or data that an expert is prepared to present at trial. All too often, however, the examination of the literature is cursory or superficial. The temptation, especially for overburdened judges not skilled in sorting through biomedical and other journals, is to check that there are articles on point, and if  the theory has been discussed (critically or otherwise) in the literature, to write that the "peer review and publication" factor supports admission of the testimony.

One area in which this dynamic is apparent is traditional testimony of firearms examiners matching marks from guns to bullets or shell casings. \2/ Defendants have strenuously objected that traditional associations of particular guns to ammunition components is an inscrutable judgment call that does not pass muster under Daubert. Perhaps the most meticulous analysis of this issue comes from an unpublished opinion of Judge Todd Edelman in United States v. Tibbs. \3/ Judge Edelman's discussion of peer review and publication is unusually thorough and may have been penned as an antidote to the strategy in which the government gives the court a laundry list of articles that have discussed the procedure and the court checks off the "peer review and publication" box.

Being an opinion for a trial court (the District of Columbia Superior Court), Tibbs is not binding precedent for that court or any other, but it has not gone unnoticed. Two federal district courts recently reached mutually opposing conclusions about Judge Edelman's analysis of one large segment of the literature cited in support of admitting match determinations--namely, the extensive research reported in the AFTE Journal. ("AFTE" stands for the Association of Firearms and Toolmark Examiners. The organization was formed in 1969 in "recognition of the need for the interchange of information, methods, development of standards, and the furtherance of research, [by] a group of skilled and ethical firearm and/or toolmark examiners" who "stand prepared to give voice to this otherwise mute evidence." \4/)

Tibbs' Analysis of the AFTE Journal

Because of the AFTE Journal's orientation and editorial process, Tibbs did not give "the sheer number of studies conducted and published" there much weight. \5/ Judge Edelman made essentially four points about the journal:

  • Contrary to the testimony of the government’s experts, post-publication comments or later articles are not normally considered to be “peer review”;
  • The AFTE pre-publication peer review process is “open,” meaning that “both the author and reviewer know the other's identity and may contact each other during the review process”;
  • The reviewers who form the editorial board are all “members of AFTE” who may well “be trained and experienced in the field of firearms and toolmark examination, but do not necessarily have any ... training in research design and methodology” and who “have a vested, career-based interest in publishing studies that validate their own field and methodologies”; and
  • “AFTE does not make this publication generally available to the public or to ... reviewers and commentators outside of the organization's membership [and] unlike other scientific journals, the AFTE Journal ... cannot even be obtained in university libraries.” \6/

The court contrasted these aspects of the journal’s peer review to a "double-blind" process and observed that the AFTE “open” process was “highly unusual for the publication of empirical scientific research.” \7/ The full opinion, which develops these ideas more completely can be found online.

Shipp

Senior Judge Nicholas Garaufis of the Eastern District of New York was impressed with "this thorough opinion." \8/ His opinion in United States v. Shipp referred to the "several pages analyzing the AFTE Journal's peer review process [that] highlight[] several reasons for assigning less weight to articles published in the AFTE Journal than in other publications" and added that

The court shares these concerns about the AFTE Journal's peer review process. In particular, the court is concerned that the reviewers, who are all members of the AFTE, have a vested, career-based interest in publishing studies that validate their own field and methodologies. Also concerning is the possibility that the reviewers may be trained and experienced in the field of firearms and toolmark identification, but [may] not necessarily have any specialized or even relevant training in research design and methodology. \9/

Harris

In contrast, Judge Rudolph Contreras of the U.S. District Court for the District of Columbia, writing in United States v. Harris, \10/ had nothing complimentary to say about Tibbs. This court defended the AFTE Journal research articles said to demonstrate the validity of firearms-toolmark identification with two rejoinders to Tibbs. First, Judge Contreras maintained that “there is far from consensus in the scientific community that double-blind peer review is the only meaningful kind of peer review.” \11/ This is true enough, but the issue raised by the criticism of “open” review is not whether double-blind review is better than single-blind review (in which the author does not know the identity of the referees) or some other system. It is whether “open” review conducted exclusively by AFTE members is the kind of peer review envisioned as a strong indicator of scientific soundness in Daubert. The factors enumerated in Tibbs make that a serious question.

Second, Judge Contreras observed that the Journal of Forensic Sciences, which uses double-blind review, republished one AFTE study. This solitary event, the Harris opinion suggests, is a “compelling” rebuttal of “the allegation by Judge Edelman in Tibbs that the AFTE Journal does not provide 'meaningful' review." \12/ But Judge Edelman never proposed that every article in the AFTE journal was without scientific merit. Rather, his point was far less extreme. It was merely that courts should not “accept at face value the assertions regarding the adequacy of the journal's peer review process.” \13/ That one article—or even dozens—published in the AFTE Journal could have been published in other journals reveals very little about the level and quality of AFTE review. After all, even a completely fraudulent review process that accepted articles for publication by flipping a coin would result in the publication of some excellent articles—but not because the review process was meaningful or trustworthy. In addition, one might ask whether the very fact that an article had to be republished in a more widely read journal fortifies the fourth point in Tibbs, that the journal’s circulation is too restricted to make its publications part of the mainstream scientific literature. The discussion of peer review and publication in Harris ignores this concern.

Beyond the AFTE Journal

The significant concerns exposed in Tibbs do not prove that the peer-reviewed scientific literature, taken as a whole, undermines firearms identification as commonly practiced. They simply mean that the list of publications over the years in the AFTE Journal may not be entitled to great weight in evaluating whether the scientific literature supports the claim of firearms and toolmark examiners to be able to supply generally accurate and reliable "opinions relative to evidence which otherwise stands mute before the bar of justice." \14/

Fortunately, newer peer-reviewed studies exist, and not all the older research appears in the AFTE Journal. \15/ Thus, the Harris court asserted that

[E]ven if the Court were to discount the numerous peer-reviewed studies published in the AFTE Journal, Mr. Weller's affidavit also cites to forty-seven other scientific studies in the field of firearm and toolmark identification that have been published in eleven other peer-reviewed scientific journals. This alone would fulfill the required publication and peer review requirement. \16/

The last sentence could be misunderstood. As a statement that the 47 studies could be the basis of an scientifically informed judgment about the validity of firearms-toolmark matching, the conclusion is correct. As a statement that checking the "peer review and publication" box on the basis of a large number of studies published in the right places "alone" is a reason to admit the challenged testimony, it would be more problematic. The "required ... requirement" (to the extent Daubert imposes one) is for a substantial body of peer-reviewed papers that form a solid foundation for a scientific assessment of a method. Unless this research literature is actually supportive of the method, however, satisfying the "the required publication and peer review requirement" is not a reason to admit the evidence. 

Do the 47 studies (old and new) in widely accessible, quality journals all show that examiners' opinions derived from comparing toolmarks are consistently correct and stable for the kinds of comparisons made in practice? If so, then it is high time to stop the arguments over scientific validity. If not, if the 47 studies are of varying quality, scope, and relevance to ascertaining how repeatable, reproducible, and accurate the opinions rendered by firearms-toolmark examiners are, then there is room for further analysis of whether and how these experts can provide valuable information for the legal factfinders.

NOTES

  1. 509 U.S. 579 (1993),
  2. No. 2016 CF1 19431, 2019 D.C. Super. LEXIS 9, 2019 WL 4359486 (D.C. Super. Ct., Sept. 5, 2019).
  3. "[T]he process that most firearms examiners use when analyzing evidence" is desctibed in graphic detail in "[t]he Firearms Process Map, which captures the ‘as-is’ state of firearms examination, provides details about the procedures, methods and decision points most frequently encountered in firearms examination." NIST, OSAC's Firearms & Toolmarks Subcommittee Develops Firearms Process Map Jan. 19, 2021, https://www.nist.gov/news-events/news/2021/01/osacs-firearms-toolmarks-subcommittee-develops-firearms-process-map.
  4. AFTE Bylaws, Preamble, https://afte.org/about-us/bylaws
  5. 2019 D.C. Super. LEXIS 9, at *35. For a decade or so, both legal academics and forensic scientists had pointed to the AFTE Journal as an example of a practitioner-oriented outlet for publications that did not follow the peer review and publication practices of other scientific journals. See, e.g., David H. Kaye, Firearm-Mark Evidence: Looking Back and Looking Ahead, 68 Case W. Res. L. Rev. 723 (2018); Jennifer L. Mnook--n et al., The Need for a Research Culture in the Forensic Sciences, 58 UCLA L. Rev. 725 (2011).
  6. 2019 D.C. Super. LEXIS 9, at *32-*33.
  7. Id. at *33.
  8. United States v. Shipp, 422 F.Supp.3d 762, 776 (E.D.N.Y. 2019).
  9. Id. (citations and internal quotation marks omitted). Nevertheless, the court found "sufficient peer review." It wrote that "even assigning limited weight to the substantial fraction of the literature that is published in the AFTE Journal, this factor still weighs in favor of admissibility. Daubert found the existence of peer-reviewed literature important because “submission to the scrutiny of the scientific community ... increases the likelihood that substantive flaws in the methodology will be detected.” Daubert, 509 U.S. at 593. Despite AFTE Journal’s open peer-review process, the AFTE Theory has still been subjected to significant scrutiny. ... Therefore, the court finds that the AFTE Theory has been sufficiently subjected to 'peer review and publication' [outside of the AFTE Journal].” Daubert, 509 U.S. at 594."
  10. 502 F.Supp.3d 28 (D.D.C. 2020).
  11. Id. at 40.
  12. Id.
  13. Tibbs, 2019 D.C. Super. LEXIS 9, at *29.
  14. AFTE Bylaws, Preamble, https://afte.org/about-us/bylaws.
  15. AFTE has sought to remedy at least one complained-of feature of its peer review process. In 2020, it instituted the double-blind peer review that the Harris court found unnecessary. AFTE Peer Review Process – January 2020, https://afte.org/afte-journal/afte-journal-peer-review-process. Whether the qualifications and backgrounds of the journal's referrees have been changed is not apparent from the AFTE website.
  16. Harris, 502 F.Supp.3d at 40.