Showing posts with label FBI. Show all posts
Showing posts with label FBI. Show all posts

Saturday, December 25, 2021

The FBI's Misinformation Campaign on Firearms-toolmark Testimony

On Tuesday (21 December 2021), the Texas Forensic Science Commission issued a Statement Regarding 'Alternate Firearms Opinion Terminology'. It is a forceful correction to misinformation from the FBI Laboratory's Assistant General Counsel, Jim Agar II. \1/ The email that attracted the Commission's critical attention tells forensic analysts what they are supposed to say in opposition to motions to limit their testimony about firearms-toolmark comparisons. As previous postings show, there has been no shortage of defense motions seeking to forbid eliciting opinions that ammunition components associated with a crime came from a particular gun.

The FBI advice to firearms examiners is entitled "Dealing with Alternate Firearms Opinion Terminology" (hereinafter Dealing). It begins by dismissing the best efforts of federal and state judges to respond to weaknesses in traditional "This is the gun!" testimony as "wholesale attempts to rewrite the firearm expert's testimony by a layman with no experience in forensic science." \2/ The fact that eminent scientists and respected jurists have questioned source-attribution testimony in general and in this field in particular does not seem to matter. According to Dealing, the limitations are "not supported by either science or the law." Despite the government's annoyance with lay judges' rulings, however, courts have a duty to review the scientific and scholarly literature to decide whether strong claims of source attributions are sufficiently warranted. \3/

Dealing continues, more reasonably, with the strategic recommendation that "firearms examiners and prosecutors should address the terminology issue head-on during their direct examination at the admissibility hearing. Preempt this issue early. Don't wait for the judge or the defense counsel to bring it up." But the tactics for bringing it up are over the top. Dealing imagines the following colloquy:

Prosecutor: Can you testify truthfully that your opinion is that the cartridge cases and/or bullets in this case
   • "Could or may have been fired by this gun?"
   • ''Are consistent with having been fired by this gun?"
   • "Are more likely than not having been fired by this gun?"
   • "Cannot be excluded as having been fired by this gun?"
Examiner: No, I cannot testify truthfully to any of those statements or just the class characteristics alone.
Prosecutor: Why not?
Examiner: For three reasons: First, there are no empirical studies or science to backup any of those statements or terminology. Second, those statements are not endorsed nor approved by my laboratory, any nationally recognized forensic science organization, law enforcement, or the Department of Justice. Third, those statements are false as they do not reflect my true opinion of identification. Such statements would mislead the jury about my opinion in this case. It would also constitute a substantive and material change to my opinion from one of Identification to Inconclusive. This would constitute perjury on my part for I would not be telling the jury the whole truth.

The "three reasons" border on the absurd (if they do not cross the border). First, the empirical studies that prosecutors cite to support the ability of firearms experts to match ammunition components to specific guns also support the bulleted statements. This is because the alternatives are lesser included statements, so to speak. If a categorical source attribution is correct, then a weaker included statement such as "cannot be excluded" also is true. If "empirical studies or science" do not adequately support these weaker statements, then, a fortiori, they do not support the much stronger claims that Dealing advocates.

Second, that law enforcement organizations and crime laboratories do not approve of the policy of replacing traditional "This is the gun!" testimony with a less telling alternative proves nothing about whether the bulleted statements are true or false. It merely means that a laboratory is unwilling to change its standard operating procedure and that "law enforcement" opposes losing the opinions that prosecutor's love their experts to provide. No self-respecting expert can say that the desire of "law enforcement" and crime laboratories for the strongest possible testimony makes less compelling testimony "untruthful."

Finally, that any lawyer -- let alone one representing the FBI -- would ask a forensic examiner to tell a judge that it would be perjurious to testify in the bulleted ways is shocking. A federal perjury prosecution would be laughed out of court. Under federal law, statements that are known to be incomplete, or, worse, fully intended to distract or mislead, do not constitute perjury if they are literally true. The leading case is Bronson v. United States. \4/ There, defendant testified as follows:

Q. Do you have any bank accounts in Swiss banks, Mr. Bronston?
A. No, sir.
Q. Have you ever?
A. The company had an account there for about six months, in Zurich.
Q. Have you any nominees who have bank accounts in Swiss banks?
A. No, sir.
Q. Have you ever?
A. No, sir.

In reality, the witness had previously maintained and had made deposits to and withdrawals from a personal bank account in Geneva, Switzerland. Clearly, his answers were calculated to avoid revealing this fact. However, the Supreme Court unanimously reversed a conviction for perjury, concluding that the federal statute did not criminalize lying by omission and misdirection.

To be sure, some state statutes define the crime to encompass wilful omissions, but the core idea remains that perjury occurs when the witness intends to give the questioner false information or a false impression so as to obstruct the ascertainment of the truth. \5/ An expert witness who testifies sincerely to true statements such as "the defendant's gun cannot be excluded as the one that fired the recovered bullet" or "measurements of the bullet and the pistol showed them both to be 9 mm, so the bullet could have been fired from the gun," is not intending to lead anyone to a false conclusion. That the FBI would like firearms examiners to give more incriminating opinions does not make the lesser included testimony false or misleading. A prosecutor who truly is worried that "[t]estimony about class characteristics alone may falsely imply an examiner was unable to reach a conclusion of identification" can ask the court to instruct the jurors that the rules of evidence no longer allow an expert witness to testify that a bullet came from a particular gun and that they may not draw any inference from the absence of such inadmissible testimony. Instead, they are to use only the testimony that the expert gave in coming to a conclusion about which gun fired the recovered bullet.

After maintaining that "laymen" (courts) are asking toolmark examiners to commit perjury, Dealing gives another specious argument to persuade toolmark experts to stick to their guns (sorry about that) and refuse to "agree to testify to the terms of 'Could or may have fired,' or 'Consistent with,' 'More likely than not,' or 'Cannot be excluded.'" FBI counsel believes that examiners who testify this way when they feel that a traditional source attribution is justified "are ratifying these bogus statements and adopting this as their testimony, giving the judge a pass on the difficult decision to admit or exclude their testimony. They are also acquiescing to the judge's faulty terminology."

This is nonsense. The law has a spectrum of options ranging from excluding every bit of information a firearms expert might provide (which is unjustified given what is known about the performance of these experts) to unfettered admission of "This is the gun!" testimony (which is traditional). The only "fault" in the intermediate testimony is that it is not as strong as a prosecutor might want it to be. It is conservative in the sense of understating probative value (as FBI counsel understands the science), but testifying conservatively at trial when that is what a court requires does not "ratify" anything about the court's ruling. It simply presents a permissible opinion. DNA experts who testified to "ceiling" probabilities of random matches because that was the best the prosecution could get some courts to accept circa 1995 were not perceived as "ratifying these bogus statements." \6/

Dealing disagrees. FBI counsel insists that "acquiescing" in court rulings is "fatal" to an examiner's career as a witness:

This is fatal. Why? Once you testify to these bogus terms, you are wedded to them for life. At subsequent trials, defense counsel will pull out the verbatim transcript of the examiner's previous testimony where they used these court-induced terms. On cross examination, they will confront the examiner with their previous testimony and contrast their opinion of "Identification" with those in previous cases, then claim the expert is merely making this stuff up. The examiner no longer has any credibility in the jury's eyes.

This fear of cross-examination is fanciful. If the expert testifies at the admissibility stage (as Dealing contemplates) that "This is the gun!" testimony is scientifically justified, then that is what the expert is on record as stating. Later, more circumscribed testimony pursuant to court order is not an inconsistent statement useful for impeachment. Any competent expert witness will have no trouble explaining that in the earlier case, I reached the conclusion of "identification" (just read my case notes), and I used other terminology only because the prosecutor asking the question (or the judge) said I had to use the lesser included language because of a legal rule rather than a scientific principle.

In contrast, the witness who follows FBI counsel's advice will lose all credibility. The truth is that the lesser included testimony, while less powerful, is no less truthful than "This is the gun!" testimony. It is somewhat like choosing a wider confidence interval to increase the coverage probability; the statement becomes less precise, but it is more likely to be true. Talk of perjury and being asked to lie suggests either that (1) the witness does not understand a statement such as "the recovered bullet could have come from/is consistent with coming from/is not excluded as coming from/is more likely to have come from the firearm in question or that (2) the witness has chosen to lobby for the prosecution rather than to educate the judge impartially.

NOTES

  1. Mr. Agar is a decorated, retired Colonel with "31 years of successful experience leading complex legal organizations as a general counsel, attorney, leader, mentor and trainer of FBI legal offices and senior-level Army staffs" and "hands-on experience in advising senior FBI and Army leaders in all legal matters." His work as Assistant General Counsel for the "FBI Forensic Laboratory" began in October 2016. On Linkedin, from which these quotations are taken, he summarizes his current position as
    Legal advisor to the largest and best forensic laboratory in the world with a staff of over 700 scientists and a budget of $110 million. Responsible for training and qualifying the FBI’s forensic examiners to testify in any and all courts nationwide and internationally, consisting of over 120 examiners in 37 different disciplines. Coordinate all discovery for the Laboratory. Provide ethics advice to Laboratory personnel.
  2. Discussion of this line of cases can be found in David H. Kaye et al., Wigmore on Evidence: Expert Evidence (3d ed. 2021).
  3. The track record of the courts in translating this literature and the growing research on firearms-toolmark comparisons into appropriate constraints on proposed expert testimony is not perfect. Indeed, most of the judicial palliatives for perceived expert overclaiming (such as the supposed limitation of "a reasonable degree of ballistic certainty" and the alternatives listed in Dealing) are far from optimal. Id. (and other postings in this blog). But these failures hardly mean that, as "laymen," judges are disqualified from trying to improve the presentation of expert knowledge by excluding certain forms of testimony.
  4. Bronston v. United States, 409 U.S. 352 (1973).
  5. See Ira P. Robbins, Perjury by Omission, 97 Wash. U. L. Rev. 265 (2019).
  6. See, e.g., David H. Kaye, The Double Helix and the Law of Evidence (2010).

Wednesday, March 20, 2019

Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 3)

Yesterday, I presented some of the testimony behind the one-in-650-biliion probability quoted in ProPublica's reporting on FBI testimony in United States v. McKreith. The figure came from the uniform probability model as applied to the placement of lines on plaid shirts. The selection of that probability model was motivated by observations of the manufacturing process.

Physical scientists are trained to make rough approximations of quantities — both large an small — and they should not be pilloried for trying to use a simple probability model to get a sense of how rare an event might be. But experts should not testify to rough calculations of extremely persuasive probabilities without any error bounds and without trying to verify the assumptions of the modeling effort by collecting and analyzing a reasonable amount of data.

The prosecution's enthusiasm for the largely theoretical probability model in McKreith extended not merely to the shirt mentioned in the ProPublica article, but also to stripes on a handbag. Dr. Vorder Bruegge's testimony on direct examination was relatively mild. He testified, without computing any probabilities, that a photograph and a handbag were "indistinguishable" with respect to "class" and "individual" features that the jurors could see for themselves:

Q. I believe we were at the Mary Kay bag, Government Exhibit 14. [W]ere you able to identify this exhibit as something that was presented to you for ... photographic comparison purposes?
A. Yes, this is the bag that was submitted to me at the FBI laboratory for comparison in this case.
Q. [W]hat observations were you able to make in terms of the class characteristics of the bag?
A. Basically, this is a handbag that has ... two dark straps. It's got a pocket on [one] side ... . It's made primarily of ... multiple panels — two side panels, two end panels, a bottom and a secondary panel that is overlapping this side panel. It's a striped bag that has these bright snaps on the end. ... [B]asically, it's ... dark silver and black stripes on the bag. ... The manufacturer of the bag is indicated by the name Mary Kay on the side.
Q. [C]an you identify ... what distinguishing features there are about the stripes?
A. Well, the stripes are evenly sized stripes. Each one is about a quarter of an inch wide, and they are alternating black and silver stripes.
Q. At the spacing, the same consistent —
A. The spacing is consistent throughout, across the bag, yes.
Q. And you indicated that it has hand straps?
A. Yes. ...
Q. [S]howing you what is marked as Government’s Exhibit 7-EE, regarding the bank robbery at SouthTrust, is this a chart that you prepared for comparison analysis? ...
A. Yes. Government exhibit V7B-EE is a chart that I prepared. ...
Q. ... [W]ere you able to identify, from looking at the bag itself, any individual characteristics, identifying characteristics, which would make this bag unique from all the other Mary Kay bags that may have come off the assembly line at sometime, using the same fabric, being the same sized bag?
A. [T]here are some small white markings on the bag ... . It's not clear to me whether it would be marker or some kind of staining on the bag that occurs at various places around the bag. [O]n the back here [are] little white marks that could be used to differentiate this bag from all of the bags.
      As far as the manufacturing characteristics, we've got another repeating pattern, much like the one in the shirt. in which we've got dark bright stripes. [I]f we look at this end panel, for example, you'll see that the very top stripe on this end of the bag is a ... totally black stripe. And if we look at this end ... where the top is, ... there's actually a little silver there. Likewise, if you look at the edges ... [at] the very top of this [side panel] is about half of one of those silver lines. On the other side, it's not quite a half of one of those silver lines. [Also, it’s] basically silver on the top of the sides — silver on the top of this side panel, end panel, but black [with] maybe a little bit of silver there, [but] when it's folded over, it's black on the top. Also, if you look at where the end panel meets the side panel, you've basically got ... the silver effectively lining up with the black, going from the end panel to the side panel. At the other end, it's a slight offset, slightly different, where it's kind of half and half. It doesn't match up exactly.
Q. ... Are the results of the randomly identifying features in terms of the way the bag was manufactured when these parts were sewn together?
A. Now, I have never been to a bag manufacturing plant, but assuming that the same sewing practices were used —
MR. HOWES: Judge, I'm going to object.
COURT: Sustained.
BY MR, STEFIN:
Q. [S]o you don't know the manufacturing process with respect to that particular bag?
A. That's correct.
Q. Okay. Were you able to do a comparison in any regard to determine whether or not there were points of identification which are similar to the government's Exhibit 14 with the bag depicted in the robbery photos from the SouthTrust bank robbery?
A. Yes I was. ...
Q. And what ... did your comparison yield?
A. Basically, I found a similarity in class characteristics between the bag, the Mary Kay bag — Government’s Exhibit 14 — and the bag carried by the Robert in the SouthTrust bank robbery as depicted on the left hand side of Government’s Exhibit VB7-EE. ...
Q. Are you able to offer an opinion as to whether Government Exhibit 14 is indistinguishable from the bag that's depicted in the bank surveillance photographs of the SouthTrust Bank robbery?
A. Yes.
Q. And what is your opinion?
A. This Government’s Exhibit 14 is indistinguishable from the bag in Government Exhibit 7-EE.

The court sustained the objection to testimony about "randomly identifying features in terms of the way the bag was manufactured when these parts were sewn together" because Dr. Vorde Bruegge forthrightly acknowledged that he was assuming certain facts not in evidence (and outside his expertise as an image analyst). Evidently, the court wanted more than a mere assumption "that the same sewing practices were used." But the next day, on re-direct examination, the prosecutor had Dr. Vorde Bruegge present the same probability model for the placement of the bag's stripes that he had used for the shirt:

Q. ... And with respect to the Mary Kay bag, Government Exhibit 14, didn't you identify individual characteristics of that bag which makes it different than other Mary Kay bags that may have come off the same assembly line?
A. Yes, I did.
Q. And in fact, how many different characteristics were you able to identify looking at that exhibit, in comparison with the bank surveillance photographs of a bag being carried by the robber?
A. There were four specific characteristics that I noted.
Q. And would you remind us of what those four individual characteristics were?
A. The first one was the alignment of the black and silver stripes from the back side of the bag with the end of the bag, the fact that the silver lines on the inside line up with the black lines on the back side. The second characteristic was the location of the snaps at the top on a silver line. The third characteristic was the very small silver line at the top of the back piece. And the last characteristic was the silver line at the very top of the back piece.
Q. And did you come up with any ... odds or probabilities that these items would appear exactly as they are on that bag in a random fashion?
A. Yes, I did. ...
Q. [H]ow were you able to arrive at a probability as far as the individual characteristic that would exist?
A. Basically I'm dealing with a black or white situation. In this case, black or silver. Either you're going to get the black line in one place or you're going to get the silver line in that place. I'm not breaking down by 50% of the black line or 50% of the silver line. I'm just saying it's either a black line or a silver line, which is a 50/50. You got like one chance in two of a specific feature being black or silver. In particular, these silver snaps on the end can either be on a silver line or a black line. They're on a silver line. That eliminates all of the other bags that would have the snaps on a black line.
      Likewise at the top, there is either a silver line at the top or there's a black line at the top. One chance in two, 50/50. So with this, the snaps and the top of the side of the back, it’s one in four. With the addition of the back of the bag silver at the top, it's 1 in 8 — 2 times 2 times 2. And then with the sides here having silver aligning with black, the silver’s either going to align with black, or the silver’s going to align with silver. That's another one in two chance. So 2 times 2 times 2 is 1 in 16.
Q. 2 times 2 times 2 times 2?
A. Yes, correct. 2 to the 4th power.
Q. 2 to the 4th power. So it is possible then to eliminate 15 out of 16 silver bags that would be coming the manufacturing process from whatever company made those bags?
A. That would be the hypothesis, correct.
Q. And did you, fact, find those four same characteristics in the photographs depicting the robber carrying the same bag?
A. Yes. Yes, I did.

This computation of 1/16 for the probability of a four-feature match sounds suspiciously like an application of the principle of insufficient reason discussed yesterday. Every feature is "black or white," present or absent. Not knowing which is more probable, we can presume that each state (present or absent) is equally probable. Four independent features then create 16 equally likely states of nature.

Over a century ago, the New York Court of Appeals soundly rejected this Laplacean reasoning. In People v. Risley, 108 N.E. 200 (N.Y. 1915), a mathematician "was permitted to testify that, by the application of the law of mathematical probabilities, the chance of such defects [in letters typed on an allegedly altered affidavit] being produced by another typewriting machine was so small as to be practically a negative quantity." The mathematics professor "defined the law of probabilities as 'a proper fraction expressing the ratio of the number of ways an event may happen, divided by the total number of ways in which it can happen.'" The court wrote that the extended multiplication of one-half for each peculiarity "was not based upon actual observed data, but was simply speculative ... ."

If the choice of 1/2 for the probability of each binary feature on the handbag was based on nothing more than the fact that that there are two possibilities -- present or absent -- then it too is "simply speculative." If it was based on the more plausible assumption that the bag is sewn together in a way that would be expected to produce a uniform distribution, then the objection -- that the expert had not even visited a Mary Kay plant to learn how the bags were made -- applies. But McKreith's lawyer did not renew the objection. Neither did he argue that the applicability of the model was not verified by data on a sample of Mary Kay bags. If anything, the 1/16 figure for the four-feature handbag match is more speculative than the one in 650 billion probability of the eight-seam shirt match.

Risley and McKreith are not the only cases in which experts have multiplied a lot of small fractions together to get a smaller number. In the 1968 California case of People v. Collins, a prosecutor had a local mathematics professor testify to the product (one in 12 million) of a series of probabilities for characteristics supplied by eyewitnesses who desctibed an interracial couple that drove a yellow automobile away from the scene of a robbery. The probabilities were data-free estimates that the prosecutor fed to the mathematician. The California Supreme Court reversed the conviction, famously remarking that "[m]athematics, a veritable sorcerer in our computerized society, while assisting the trier of fact in the search for truth, must not cast a spell over him," and sparking much distrust in the legal community of probability calculations.

However, the probability model in McKreith, for handbags as well as shirts, is more plausible than the one in Collins. For one thing, the assumption of uncorrelated features is more plausible, and there is at least a modicum of knowledge of the process generating the features. Nevertheless, as I observed with regard to another case in which appellate lawyers for a defendant put an explicitly probabilistic assessment of evidence into their brief, "[t]he attempt to use probability theory ... was heroic. Like many acts of heroism, it also was hasty. Although there were some measurements and estimates of quantities bearing on guilt or innocence, the empirical data were so sketchy that the computations inevitably were more creative than convincing." 1/

NOTES
  1. D. H. Kaye, Book Review, Statistics for Lawyers and Law for Statistics, 89 Mich. L. Rev. 1520, 1543 (1991).
POSTINGS IN THIS SERIES
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 1), Mar. 3, 2019
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 2), Mar. 19, 2019.
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 3), Mar. 20, 2019

Tuesday, March 19, 2019

Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 2)

As noted earlier this month, one of the more striking parts of the ProPublica articles on the FBI laboratory's image image analysis unit 1/ is its discovery of testimony, given more than sixteen-years ago, about the probability of finding certain features in photos of a shirt worn by a bank robber caught on film in a string of robberies. The government accused Wilbert McKreith of eight robberies and of and possession of firearms having been previously convicted of a felony. 2/

According to ProPublica, a ruling that Dr. Richard Vorder Bruegge’s “testimony met the Daubert standard” “enshrined the FBI unit’s techniques and testimony as reliable scientific evidence.” This characterization seems to overstate the importance of a single, unpublished pretrial ruling that followed a short, spur-of-the-moment hearing. 2/ Still, the testimony was central to McKreith’s conviction on the bank robbery counts. Dr. Vorder Bruegge’s conclusions fortified the testimony from ordinary witnesses that, among other things, McKreith and the bank robber wore two wrist watches at the same time; wore or had ski masks in south Florida; wore plaid shirts; drove a maroon, burgundy or red-colored car; and had the same general appearance.

Dr. Vorder Bruegge’s signal contribution came from comparing video images from bank cameras to items seized from the defendant’s home and to McKreith himself. Some of this testimony was not definitive. For example, Vorder Bruegge testified that enlargements of the photographs from several bank robberies lacked sufficient resolution for him to say “with a scientific certainty” what that they showed wristwatches. Instead, he stated that they were “consistent with these features being two watches on the left wrist of the bank robber.” Likewise, he testified that a bank photograph that showed a profile of the robber’s face lacked the resolution to reveal allegedly “individual identifying characteristics [such as moles, stars, chipped teeth, ear patterns or other facial minutiae],” but “the overall characteristics of the profile, which include the shape of the nose, mouse, and chin” displayed “similarities” “consistent with” McKreith’s being the bank robber.

Dr. Vorde Bruegge’s analysis of the pattern of a plaid Van Heusen shirt in photos from seven of the robberies produced more conclusive evidence. A condensed version of the testimony on direct examination on December 16, 2002, is in the shaded boxes that follow. Discussion is interspersed between the boxes. ProPublica posted the full transcript.of the testimony from Dec. 16 and 17.

"Individual Identifying Characteristics"

MR. [Roger] STEFIN [Assistant US Attorney]: Our next witness would be Richard Vorder Bruegge.
THE COURT: Okay, members of the jury, [o]ur next witness will be an expert witness. ....
[Defense counsel objected that the government first needed “to lay a sufficient predicate” and “request[ed] a Daubert ... hearing. The court held an impromptu hearing, with no written briefing, at which Dr. Vorder Bruegge affirmed that “the techniques [are] well recognized ... in forensics [a]nd in the scientific community at large.” The court overruled the defendant’s objection.]
Q. [by MR. STEPHIN] [H]ave you had the opportunity to view the eight by ten photographs that have been introduced in evidence in this particular case with respect to the bank robbery images in each of the eight bank robberies?
A. Yes, I have. ...
Q. ... Let’s talk about the shirt now. What is it about the manufacturing process of shirts that might enable you to identify features of that shirt that could be then compared with image analyses?
A. Well, any, any comparison analysis involves a comparison of first of all, the class characteristics. ... In a shirt like this, the class characteristics ... include such things as, does it have a pattern, which this shirt does — it has a plaid pattern or a checked pattern, if you will. ... It also is a button-down shirt — it's not a pull-over, and it's a long-sleeved shirt. ...
Once one has ... found ... class characteristics that match, one can move on to look at the individual identifying characteristics. ... [I]f I were trying to differentiate one person from another and reach a positive identification, then I would need individual identifying characteristics, such as moles, scars, freckle patterns, chipped teeth. These, these features enable you to really differentiate people down to saying this person is unique from all other people. [B]ecause this shirt is a patterned shirt, there are individual identifying characteristics ... based upon ... the way ... the pieces on the shirt are cut out and then sewn together.

The class-individual distinction is entrenched in forensic science, and there is a grain of truth in it, But it assumes what is to be proved -- that every "individual characteristic" (or some unspecified combination of them) makes an individual distinguishable from every other individual. "Class characteristics" are known to be generic, whereas "individual" ones might not be. Logically, both work the same way -- the presence or absence of a characteristic changes the probability of a common source for the specimens or images being compared. A less tendentious pair of terms would be "generic" and "randomly acquired or incorporated," keeping in mind that random characteristics are not necessarily specific to an individual.

"One in-35 Through Random Processes"

Q. So what can you conclude with respect to the pattern itself on this shirt?
A. [B]ecause we don't see the patterns matching across the seams, we can use that as a way to individualize this shirt relative to the other shirts that would be manufactured at the same time. ...
Q. [W]hat do you call this when you find an individual feature?
A. Individual identify characteristic. Basically, each seam ... can be considered on its own as an entire set of individual identifying characteristics. It's almost one individual identifying characteristic. ... [I]f we look at where the right yoke meets the right front panel, there's this ... thick dark that I've been talking about this comes down from the neck, and it terminates here about 2/3 of the way across this panel ... . Likewise, if we look at the left sleeve, I point out that ... very thin dark line here, which is coming up the right sleeve. It almost exactly meets the point where the yoke joins the front panel. ... Likewise, ... look at the collar itself. You'll see that the collar has a very dark line [that] just cuts the corner of the bottom hole on this side. That is paralleled on the other side because it's one piece. So that's another individuating characteristic for this shirt. ...
Q. All right. And did you take precise measurements of these types of lining up of the different — the thick lines or the thin lines in order to further identify, you know, the individual characteristics of this particular shirt?
A. I measured the width of these features on this shirt so that I could figure out [the] relationship ... between places meeting on one side with those meeting on the other side.
Q. And I think you said that the pattern repeats basically every three-and-one-half inches?
A. Yes. Basically, ... if we just repeat going from this dark line to the next dark line, it's three-and-a-half inches. ...
Q. The thick line or the thin line?
A. The thin line. Each ... of these features repeats every three-and-a-half inches. It's just that the thin line is the smallest feature that we can see very easily.
Q. So you were using the thin line as a point of reference in measuring out the three-and-one-half inches? A. Exactly. Recall that I said there's a curved surfaces here where the sleeve meets the other seams. That would complicate any type of repeat analysis because since it's not a straight line, that's going perpendicular to that feature, so it's actually going to be longer.
So basically, if we did have a single seam that we were trying to match this pattern to run across it, the chances of that very thin dark line matching up, matching up or at least touching the black line on the other side, would be one in 35 through random processes.
Q. Could you explain that, explain that a little bit, how you came up to the one in 35? ...
A. ... Let's use, let’s use the yoke and the sleeve. ...
Q. OK ... And the yoke, again, is this back panel ... ?
A. Right. I'm using the yoke in the back because they're almost aligned. ... You can see here that on the yoke the dark line is ... slightly below the seam here. And you have to go about half an inch down before you hit the black line on the other side.
Q. Okay. In other words, you can measure the distance between the black line on the yoke to the black line on the sleeve, and you measured, say ... a half an inch?
A. Yeah.
Q. All right. And so just taking that one point of reference, what would be the odds that two shirts coming out the same manufacturing plant being manufactured in this, in this method would end up having the same alignment of yoke and the sleeve, whereby you could have this ... half an inch offset between them just for this one point of identification?
A. Just for that one point, it would be a 1 in 35 chance.
Q. And how did you come up with the number that it's a 1 in 35 of probability that these two items would line up the same in more than one shirt?
A. Because the feature itself that I’m looking to align is 1/35th of the overall repeat length. And to see ... that one feature come up randomly happens to be one in 35 times.
Q. So one in every 35 shirts manufactured by the company using this patterned cloth, the exact same cloth, you could say that one in 35 probability-wise would come up with the exact same alignment between the yoke and the left sleeve?
A. That is correct.

Several things are going on here to give rise to a uniform probability distribution. First, the plaid design of the shirt comes from a repeated 3.5"-wide distinctive pattern that contains lines of at least two different thicknesses. Second, these lines can be offset from one another across the seams of the shirt. Third, Dr. Vorder Bruegge measures how large the offset is in 1/10th" strips on the 8x10" photographs. Finally, the exact offset -- and hence, which strip a corresponding line on the other side of a seam falls, is determined at random.Thus, the number (call it X) of 1/10th" strips that separate the starting point of every plaid block on the two different cuts of fabric that are sewn together along a seam can take on the value x = 0, 1, 2, ..., 34, with probability f(x) = 1/35 for every x. The following sketch of a repeating block showing only one line in the pattern may clarify what x stands for:
1/10" strip ---------| |
1/10" strip          | |
1/10" strip          | |---------} x=2 (2/10" offset)
1/10" strip          | |
...
1/10" strip ---------| |
1/10" strip          | |
1/10" strip          | |---------
1/10" strip          | |
...                   s
1/10" strip ---------|e|
1/10" strip          |a|
1/10" strip          |m|--------
The idea of assigning an equal probability to elementary events dates back to Bernouilli (1713) and Laplace (1814). The underlying principle of insufficient reason, indifference, or symmetry holds that if there is no reason to believe that any event is more likely to occur than any other event in a set of possible, mutually exhaustive events, then one should assume that all the events are equally probable. This principle offers one way to motivate or understand the axioms of probability theory. Although it has fallen out of favor, it is not without defenders. 4/

But Vorder Bruegge did not base his probability model on a metaphysical or philosophical principle. He gave an empirical justification. During the short Daubert inquiry, he testified that he visited "manufacturing plants, factories where articles of clothing [are] made" and that "in this particular instance, I visited ... cutting plants in Alabama where the patterned material is cut out, as well as manufacturing plants where the cut out pieces are sewn together so that I can see for myself how the process takes place." Furthermore, "I've also been to another plant, Guess plant in Southern California, where shirts were also manufactured, and found that they use the same manufacturing process at the Guess Factory as they do at the Arrow shirt factories in Alabama and ... Georgia." He learned that "[m]anufacturers do not make an effort, in general, to make these features align, because to do so would be prohibitively expensive" and that "there are also places where it is not possible ... to make them align because of the curvature such as along the arms and the sleeves."

In essence, he proposed that busy workers stitched the pre-cut pieces of fabric together without regard to how well their patterns aligned with one another. He elaborated:

Q. All right. Could you ... explain the manufacturing process ... as to how shirts of this nature would be made in a factory?
A. [Y]ou start off with a huge bolt of cloth that can be hundreds and hundreds of yards long. The cloth is then laid down at a cutting plant on a table that [is] maybe about a hundred yards long. Now, these huge bolts will be rolled out on the table and then rolled back onto itself until the entire roll is done. The another role is attached, ... and it continues until you have on the order of 500 plies of material. Now, because of the way that the material is rolled back in, this pattern is not going to line up from one ply to the next. ...
Once the fabric is laid out and you've got 500 plies on top, the manufacturers lay out, basically, tracing paper that has cutting patterns. Just like if you were sewing at home and making your own clothes, you would have a pattern to cut out. Only this is a hundred yards long [and] has been designed by engineers whose only job is to figure out how best to place each piece of this shirt in its closest proximity so they are wasting as little of that fabric as possible. ...
[T]hey slap it down and they will actually have people get on there with jigsaws and cut them out by hand. Or in some of the plants they will have computers that can manage the cutting out. Once all of those pieces are cut out, they [are] transferred over to the sewers ... [Y]ou've got people who do nothing all day but sew on sleeves onto shoulders or yolks on to the back.
Q. ... So there will be, like, thousands of pieces for the collar and thousands ... for the sleeves and so forth?
A. ... If you have 500 plies thick, you have 500 pieces in one particular area. ... [O]n this shirt we've got two pieces on the collar. If you look closely at this shirt, you'll see there's actually a piece here it appears on the outside and a piece on the inside. There's another piece that goes around the collar; so that's three pieces. ... [All together, there are] 16 pieces of this same pattern cloth on this shirt [that must be sewn together].
Q. What significance does that have in terms of determining whether or not, you know, identifying the uniqueness of a particular shirt as opposed to any of the other shirts that are manufactured by the plant using the same cloth coloring?
A. Well, as I mentioned before, every one of these pieces is cut out. The plies line up on every row. If, in the manufacturing process, they happened to take a piece from the top layer and try to stitch it to a piece from the next layer down, you're going to get a lot of randomness in this process. Because there is this offset, you're not going to get this pattern lining up across the seam in any consistent way from one shirt to the next. ... [I]n fact, you can't get a curved seam like this to have an alignment across it, because it isn't geometrically possible.
You look at the pockets on this shirt and you'll see that they line up. [T]he fact that one of these pockets lines up makes this shirt twice as expensive to manufacture as it would be if this pocket were on the bias. [T]hey have to make sure that this pocket is cut out and exactly the same orientation as the front panel. Furthermore, they have to actually have a human being take the time to physically line this pocket up when they're sewing ... .
Q. Is there any effort to line up any of the other pieces of cloth? In other words, lining up the lines from the back to the sleeves, or the yolk to the sleeves, or the yolk to the front panels —
A. No. No. And you can see that for yourself just by comparing the way the left sleeve doesn't align with the yoke in the same way that the right sleeve aligns with the yoke. ...

The uniform-probability model of the placement of the lines is not as "preposterous" and "outrageous" as the ProPublica article suggests. But no one can see for themselves that there is no "effort to line up any of the other pieces of cloth" from the mere fact that the plaid patterns are displaced where a "sleeve aligns with the yoke." That is like saying that a marksman is shooting entirely at random because he missed the bullseye. There is a random component, but if the the marksman is skilled, bullets are more likely to arrive near the center of the target than to be dispersed equally across it.

Our theory about the marksman firing at random would be more credible if we saw that he was wearing a blindfold and firing rapidly. Likewise, the theory in McKreith is that the workers won't "take the time to physically line [the plaid patterns] up when they are sewing." It could be a good theory, but how has it been validated? Might not some workers try to get the unit patterns to line up just a little bit as they join the pre-cut pieces of fabric? If that happened, smaller cross-seam offsets would be more probable than larger ones. We could test the theory empirically by inspecting shirts. If the uniform probability model is correct, we would expect to find pretty much the same number of 1/10" offsets (x = 0, 1, 2, and so on) at a given seam in a large sample of Van Heusen shirts with the same design. The FBI apparently had no such data to support the probability model.

To be sure, the probability model was empirically motivated — Dr. Vorde Bruegge informed himself about the manufacturing process by visiting several manufacturing plants. His sense of the process might be entirely correct. When interviewed, I was not told what model he had adopted and what the basis for it was. `But even if I had been asked to study the testimony before reacting, I still might have said that even a plausible model could be "terribly flawed" (or at least not well validated).

"650 Billion, Give or Take a Few Billion"

A characteristic that results from the manufacturing process that occurs with a probability of 1/35 does not "individualize" an item in the only-one-in-the-universe sense that criminalists use the term. It is a generic feature, and Dr. Vorder Bruegge did not claim otherwise. To arrive at the ultimate opinion took a few more steps. The first was to assume that the offset at each seam is statistically independent of the offset at every other seam (and combination of them).

Q. Now, would the same randomness apply to all the other features in pieces of cloth that go into the shirt?
A. Yes, they would.
Q. So if you were able to, for example, reach a measurement as far as the left sleeve is concerned, with the yoke, would the line up — would the line up be the same or would it be, again, random with respect to the right sleeve?
A. There's going to be a 1 in 35 chance that it's going to be the same on the right sleeve as it is on the left sleeve.
Q. All right. So if you were able to find, for example, from the photographs, two points of identification, whereby the photograph matches the shirt in one area, maybe the yoke to the sleeve on the left side, and then you're able to find a second point of comparison or identification, say on the right sleeve and the yoke, what would be the odds of two shirts randomly being manufactured coming from the factory that would match this particular shirt?
A. It would be ... 1 in 35. But to simplify things and to be conservative, I prefer to use one in 30. By saying one in 30, that's — each giving it a better chance of being the same, but it makes the math easier. Thirty times 30 is 900. So one in 900 chance that you're going to find another shirt that has the left sleeve aligned to the yoke the same way and the right sleeve aligned to the yoke in the same way.
Q. All right. Now let's say you ... have a good enough picture that you can make three points identification. ...
A. Well, ... if we had sleeve to yoke, yoke to back, and yoke to sleeve, then that’s 30 times 30 times 30, which is one in 27,000.
Q. So if you were able to make three items of — points of identification the odds would be one in 27,000 that there would be two shirts randomly made from the factory that would have those three points of identification exactly the same?
A. Correct. ...
Q. And were you able to actually make those types of identification with respect to the photographs depicted in the bank robbery surveillance photos with this shirt ... .?
A. Yes. I was.
Q. And what was the highest number of points of identification that you were able to match up with respect to any particular bank robbery photo — in the surveillance photographs with respect to this shirt?
A. Eight. ...
Q. So 30 to the eighth power [would] be the odds in which two shirts would be randomly manufactured by the company [with] all those eight points of identification lining up exactly the same?
A. That's correct.
Q. And does your pocket calculator actually print out all the numbers that would come out if you were to insert 30 to the 8th power?
A. No. It came out to be 6.5 x 10 to the 11th, which is basically 650 billion, give or take a few billion.

Some of ProPublica's criticism of this part of the testimony misses the mark. The one example of the "[m]any problems in the examiner’s testimony [that] went unnoticed, or were simply unknown, during trial" is that "Vorder Bruegge undercut the precision of his calculations when he admitted having rounded down the shirt measurements used in his calculations because 'it makes the math easier.'" But how could anyone not notice that he used the figure of 1/30 instead of 1/35? And why is that reduction in "the precision of the calculations" a problem for the defendant? It means that the joint probability is even smaller than the pocket calculator's output.

When ProPublica contacted me and quoted the 1/650,000,000,000 figure, my reaction was "How could you get that number"? Not knowing anything about the case, I assumed it was the product of frequency estimates for different kinds of characteristics, such as shirt size, color, style, pattern, imperfections, discolorations, and so on. I doubted that such frequencies were known with sufficient accuracy to justify giving a single  astronomically large denominator to a jury. 5/

Apparently, "Karen Kafadar, chair of the statistics department at the University of Virginia," had a similar reaction and was among the "seven statisticians and independent forensic scientists [who] told ProPublica that "[t]he statistics were also preposterous" because "[t]he features Vorder Bruegge matched might be common in plaid shirts, making them of little value for identifying the garments." Indeed, Dr. Kafadar inveighed "that the 1-in-650-billion claim 'makes about as much sense as the statement two plus two equals five.'"

However, while the proposition that "two plus two equals five" seems to violate mathematical logic, it is not illogical to argue for a uniform probability distribution of offsets and to multiply probabilities as Dr. Vorder Bruegge did. If the 1/10" measurements are all correct, if the offsets are uniformly distributed over such strips, and if the independence assumption for different seams holds, then the probability of an equal offset at every one of n corresponding seams is (1/35)n, just as he testified. The assumptions are part of a perfectly logical argument, and they are not inherently "preposterous." But they have not been the subject of any systematic study (that I know of).

The Probability of Uniqueness

With all that said, the testimony had an unusual virtue over the typical "individualization" thinking of its day. Dr. Vorde Bruegge followed Lord Kelvin's dictum that "when you can measure what you are speaking about, and express it in numbers, you know something about it; but when you cannot measure it, when you cannot express it in numbers, your knowledge is of a meagre and unsatisfactory kind; it may be the beginning of knowledge, but you have scarcely, in your thoughts, advanced to the stage of science, whatever the matter may be." Dr. Vorde Bruegge was explicit and numerical about the grounds for concluding that only one shirt was involved in the pictures:

Q. [D]id you, during your research, contact the ... Van Heusen Company or the factory to determine ... approximately how many shirts using this cloth and this design were made by them?
A. They do not keep records of every single shirt that they ever make, but they have people who recognize patterns and they also know what their typical runs are. [A]t most, they said they would have made about no more than 18,000 of these shirts. ...
Q. So seven points of identification with respect to the bank robbery surveillance photos and the shirt with the Commerce Union bank robbery?
A. Yes.
Q. Let's go to the next robbery. ...
A. That would be seven, I believe.
Q. So the odds of this being random with two shirts, 30 to the seventh power?
A. If we use 30.
Q. Being conservative?
A. Yes. ...
A. That's seven points of identification with respect to the Bank of America shirt?
A. Yes.
Q. [F]rom the ... Bank United robbery? ... Five positive points of comparison?
A. Right.
Q. Okay. South Trust ... Well it's either 30 to the 7th power or 30 to the 8th power or 30 to the 6th power? A. That's it for South Trust. ...
Q. So that would be five points of identification for the Union Bank photographs?
A. Yes.
Q. [I]s there a point that you were able to conclude that all the photographs in all the different bank robberies are depicting the same shirt?
A. Well, for me in this case, it only took three points. Because with one to the 30th chance each time ... having three points ... I've got 30 times 30 times 30, which is 27,000, which is half again as many as 18,000. [S]o for my opinion it's enough to go over to say that three of those are enough. The fact that I've got four, five, six, seven and eight, just makes me all the more certain.
Q. And what is your opinion with respect to this comparison analysis?
A. They're all the same shirt.
Q. All right. And all the shirts. In the bank robbery surveillance photographs, do you have an opinion as to whether or not they are the same shirt as Exhibit 11, which is the questioned shirt?
A. Government Exhibit 11, in my opinion, is the shirt worn by the bank robber in each of these seven bank robberies. ...

Right or wrong, the testimony is transparent about the threshold for concluding that a picture has the defendant's shirt in it. For 18,000 shirts, Dr. Vorder Bruegge asserted, it takes only three seams to conclude that there is but one shirt in existence. But is this a reasonable criterion for an inference of uniqueness?

It might seem that way. One might be tempted to reason as follows: The random-match probability for two seams is 1/900. For 18,000 shirts manufactured as described in the uniform probability model, we would expect to find about 18,000 shirts x 1/900 matches per shirt = 200 shirts with two matching seams. That is way too many to claim individuality. But for 3 matching seams, the expected number is 18,000 x 1/27,000 = 2/3. In other words, we would expect to find lots of two-seam matches, but not even one three-seam match. So it does not look like a second shirt is very likely for three or more matching seams.
But there is a problem here that did go unrecognized in the case and the article about it. I have called it the expected value fallacy. 6/ It consists of thinking that as long as the expected number of items in a population is less than 1, the probability that there really is less than one is very high. Life, or at least mathematics, is not this simple. Even when the expected number is a little less than 1, as it is here, the probability of at least one additional matching item is appreciable. In this case, it is about 3/10, as shown in Box 1.
BOX 1. THE PROBABILITY OF ANOTHER MATCHING SHIRT
The number y of items with a given characteristic that has a constant, small probability p of appearing in each item in a large population of size n is a Poisson random variable with the parameter λ = np. The probability of any y is f(y; λ) = λye-λ/y! For the three-seam match, λ = 2/3.The conditional probability that at least one more item in the population has the characteristic, given the known fact that one such item is present is [1 – f(0, λ) – f(1; λ)] / [1 – f(0; λ)] = 0.30.

In short, even ignoring any modeling and measurement uncertainty in McKreith, there is a 30% probability that another Van Heusen shirt that matches at three seams has been manufactured. This large a risk of an erroneous conclusion of individuality would not be acceptable under a 2000 FBI policy that uses similar reasoning in explaining when an analyst may testify that a DNA profile comes from a named individual.

Thus, it appears that Dr. Vorde Bruegge choose a relatively undemanding quantitative threshold for source attribution. We can say this only because he was unusually clear as to the probabilistic basis for his conclusion that only one shirt —  the defendant's —  appeared in all the pictures. It also should be noted that the final opinion would have been the same had he selected a threshold as high as five, for which the conditional probability of matching Van Heusen shirts of the same type would be only 0.0004, or 0.04%. Still, the prosecution introduced the three-seam testimony to make the matches at five and more seams appear more impressive than they were. Even if this abuse of statistical reasoning did not affect the outcome, it was unfortunate.

NOTES
  1. Ryan Gabrielson, The FBI Says Its Photo Analysis Is Scientific Evidence. Scientists Disagree, Propublica, Jan. 17, 2019; Ryan Gabrielson, FBI Scientist’s Statements Linked Defendants to Crimes, Even When His Lab Results Didn’t, Propublica, Feb. 22, 2019.
  2. United States v. McKreith, 140 Fed.Appx. 112, 2005 WL 1600471 (11th Cir. 2005) (per curiam).
  3. Hans-Werner Sinn, A Rehabilitation of the Principle of Insufficient Reason, 94 Q. J. Econ. 493 (1980); Jon Williamson, Justifying the Principle of Indifference, 8 European J. Phil. Sci. 559 (2018).
  4. An oral ruling in the midst of a trial rarely creates or enshrines a rule of law. The U.S. Court of Appeals for the Eleventh Circuit affirmed the conviction, but it did not consider its opinion important enough to release for publication, and McKreith did not raise the Daubert claim on appeal. Although the January article states that McKreith “exhausted his appeals, most of which attempted to dispute the FBI Lab findings,” the Westlaw database’s history of the case displays only one direct appeal and one petition for postconviction relief for ineffective assistance of counsel. In neither of these attacks did McKreith challenge the scientific basis of the FBI lab’s work, and no court seems to have cited the case to support admitting similar testimony.
  5. The article states that "The statisticians who reviewed Vorder Bruegge’s materials for ProPublica said the examiner’s calculations cannot be correct. Vorder Bruegge’s statistic — 1 in 650 billion — is simply too astronomical to be true, said Kaye, the Penn State professor. There isn’t a database documenting features on plaid-shirt seams like there is for human DNA, making it impossible to determine the likelihood a different shirt would appear to match the robber’s shirt." I was not one of "[t]he statisticians who reviewed Vorder Bruegge’s materials for ProPublica," but I was (and am) of the view that estimating such small numbers on the basis of strong modeling assumptions alone is fraught with danger. The same thing can be said about DNA evidence. Decades ago, England's leading forensic statistician, Ian Evett, rhetorically asked me why American experts give such small probabilities in DNA matches instead of stopping at a number like one in a million.
  6. David H. Kaye, The Expected Value Fallacy in State v. Wright, 51 Jurimetrics J. 1 (2011).
POSTINGS IN THIS SERIES
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 1), Mar. 3, 2019
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 2), Mar. 19, 2019.
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 3), Mar. 20, 2019

Sunday, March 3, 2019

Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 1)

The FBI’s crime laboratory has had its ups and downs. Mistakes or misconduct by hair analysts, latent fingerprint examiners, explosives experts, and DNA analysts have been the subject of many a headline. 1/ In contrast, the more mundane efforts of photographic experts have not prompted extensive reports of any problems, let alone scandals. These experts testify about such matters as whether a surveillance photograph matches a photograph of a given person, 2/ the height of individuals as inferred from photographic images, 3/ and whether images in child pornography cases are those “of real, non-virtual people.” 4/

These halcyon days came to end this year when Propublica published a pair of detailed articles accusing the FBI's Forensic Audio/Video and Image Analysis Unit of relying on “flawed methods” to generate biased findings that did not deserve to be called “scientific evidence.” 5/ In an interview for the first story, the reporter told me that he discovered studies showing that analysts were unable to consistently pick out the same features from pictures and that, in one case, an FBI expert had testified that the probability of certain detailed features appearing on pictures of two different shirts would be 1 in 650 billion. He wanted to know what a statistician would say about these matters. I responded that these findings were definitely problematic, and my immediate reactions appeared in the first article:
If examiners cannot mark the same features each time they use a technique, “then you can’t rely on the result, I think that’s what any statistician would say,” said David Kaye, a Penn State University law professor and expert on DNA analysis. “It’s not a reliable measure.”
Vorder Bruegge’s statistic — 1 in 650 billion — is simply too astronomical to be true, said Kaye, the Penn State professor. There isn’t a database documenting features on plaid-shirt seams like there is for human DNA, making it impossible to determine the likelihood a different shirt would appear to match the robber’s shirt. ... “It may be an honest belief,” Kaye said, “though terribly flawed.”
At the time, I had not seen the studies and testimony in question. Having begun to read through them, I am ready to share some more informed (but still tentative) impressions.

I. Reliability

A reliable measuring instrument gives approximately the same results when used repeatedly on the same material under the same conditions. A radar gun that registers the same speed (within, say, 1 mile per hour) for all cars moving at one speed gives reliable measurements at that speed. The measurements may not be correct, but they are consistent, which is all that reliability connotes in statistics.

In pattern-matching tasks performed by human observers, the situation is slightly more complicated. If the patterns are chock full of different features that are very specific to the source of the pattern, all examiners, or even the same examiner repeating a comparison at a later time, might not use the same features in coming to a subjective judgment that the source of two patterns is the same or different. For example, one fingerprint examiner might use minutiae on one part of a print one time and minutiae from another region at other times. That kind of variation would not necessarily make the binary judgments (of same-source and different-source pairs) unreliable or invalid.

The articles cited in the Propublica exposé are not really directed at such binary judgments. The more substantial article cited as proof of unreliability proposes a more standardized and quantitative approach -- a "count-based method" -- to comparing images of hands in photographs, 6/ but it does not dismiss the current, trust-the-examiner's-global-judgment approach as invalid. Rather it states that
While there is a lack of standardization in how features are documented, [this] does not invalidate the use of such features in a photographic comparison. Future study is warranted to examine how successful examiners are when tested with comparisons leading to a conclusion in regard to individualization.
In other words, even if one or two studies indicate that image examiners are not consistent in picking out features in photographs of hands for later comparisons, it does not follow that examiners are unable to reliably and accurately classify pairs of images with respect to hands they depict. The problem -- and it is a grave problem -- is that we lack the studies to know when examiners can do what they claim (at known levels of accuracy) and when they cannot.

A reader of the Propublica articles might conclude that research has demonstrated that the FBI procedures are bunk or junk science -- the article speaks of their having been "debunked." The more accurate conclusion is that they are not science (although digital technology is part of them). Some of the underlying ideas are plausible, but their implementation has yet to be validated (or invalidated) scientifically. 7/ Looking at the write-ups of the two studies, I cannot say that they move the ball greatly..And neither of them claim to. 8/

II. Probability and Individualization

The Propublica articles singled out a leading forensic scientist — Richard Vorder Bruegge, who also chairs the Digital-Multimedia Committee of the Organization of Scientific Area Committees for Forensic Sciences (OSAC) — as a source and savior of unscientific testimony about individuals and clothing captured on film. A centerpiece of the first article is his testimony in a sixteen-year-old case about the probability of finding certain features in photos of a shirt worn by a bank robber caught on film in eight robberies.

Having spent much of my career puzzling over probability computations presented in court, I wanted to read the actual transcript. For readers who share my curiosity about how probability evidence reaches juries but who may not have extra hours on their schedules, I have greatly condensed the testimony and will present some of it together with annotations in a later posting. 9/

NOTES
  1. E.g., Mark Hansen, Crime Labs Under the Microscope After a String of Shoddy, Suspect and Fraudulent Results, ABAJ, Sept. 2013; John Kelly & Phillip Wearne, Tainting Evidence: Inside The Scandals at the FBI Crime Lab (1998); Office of the Inspector General, The FBI DNA Laboratory: A Review of Protocol and Practice Vulnerabilities, May 2004.
  2. United States v. Martin, No. Crim. 98–178, 2000 WL 233217 (E.D. Pa. Feb. 25, 2000), aff’d, 46 Fed.Appx. 119 (3d Cir. 2002) (on the basis of many similarities, Dr. Richard Vorder Bruegge came “very close to making a positive identification”).
  3. United States v. Kyler, 2015 WL 13450483, Case No.: 4:09cr30/RH/GRJ, Case No.: 4:12cv438/RH/GRJ (N.D. Fla. Feb. 26, 2015), aff’d, 429 Fed.Appx. 828 (11th Cir. 2011) (“Dr. Vorder Bruegge testified that he used reverse projection photogrammetry to analyze video images [from three robberies to] determine[] that the ... true height of the individual in the three videos was slightly above that range [of 5'10.5" and 5'11"].
  4. United States v. Rodriguez-Pacheco, 475 F.3d 434 (1st Cir. 2007).
  5. Ryan Gabrielson, The FBI Says Its Photo Analysis Is Scientific Evidence. Scientists Disagree, Propublica, Jan. 17, 2019; Ryan Gabrielson, FBI Scientist’s Statements Linked Defendants to Crimes, Even When His Lab Results Didn’t, Propublica, Feb. 22, 2019..
  6. Christina A. Malone, Michael J. Salyards & Meredith Hein, Inter-/Intra-observer Reliability of Hand Assessment Using Skin Detail: A Count-based Method, 60 J. Forensic Sci. 1605 (2015). In many perceptual tasks, the clearer the stimuli are, the more consistent the responses will be. Thus, this paper also suggests that some hand images produce more repeatable and reproducible observations of features than others. If image examiners are reasonably reliable under some conditions, it is not appropriate to dismiss all comparisons as unreliable. But rescuing some of them requires a metric for distinguishing the cases likely to lead to reliable measurements of individual features from the cases that are likely to be generate unreliable measurements.
  7. To say, as the article quoted above does, that future study is "warranted" if "individualization" is to be introduced as scientific evidence, is an understatement. Scientific evidence requires scientific validation, which translates into serious, blind testing in this context. Of course, there is no compelling reason for "individualization." Examiners can testify to their sense of the probabilities of the observed features under alternative hypotheses about the origin of those features without assessing the probabilities of the hypotheses themselves.
  8. The other study cited in the article is unpublished. It was the subject of a talk at last year's annual meeting of the American Association of Forensic Sciences. Derek A. Boyd, Aislynn MacKenzie, Briana M. Turner-Gilmore, Richard Vorder Bruegge et al., Observer Agreement in the Identification and Quantification of Dorsal Hand Traits From Digital Images (abstract). The abstract acknowledges the limited proof of the validity of comparisons of "dorsal hand traits in the identification of perpetrators and victims of these criminal activities through photographic comparison." In fact, the FBI researchers state that "[t]he qualitative nature of this method prevents it from meeting Daubert standards, as there are no known error rates associated with rates of identification of these traits from visual media." Their study of the back of one hand and six examiners checked for reliability in counting "scars, moles, freckles, and knuckle skin-creases." The results were mixed:
    Calculated coefficients of variance indicated high levels of data dispersion among scars (cv=1.206), moles (cv=1.546), and freckles (cv=1.270). Coefficients of variance calculated for counts of knuckle skin-creases on each digit suggested comparatively lower levels of dispersion (cv1=0.419; cv2=0.404; cv3=0.450; cv4=0.530; cv5=0.354). Tests of intra-observer error indicated a statistically significant difference in mean counts between first and second observations of freckles (t=-2.43, df=11, p=0.034) and knuckle skin-creases on the second digit (t=-2.80, df=11, p=0.017), but not for any other traits observed.
    Testing the hypothesis that there is no difference at all in mean counts is an odd way to discuss reliability, but it sounds like the researchers were trying to say that only a few features were not reliably measured. At least, they wrote that "[t]his exploratory study found that most traits exhibited statistically minimal intra- and inter-observe disagreement" (emphasis added). But then they wrote the sentence that Propublica highlighted: "There is considerable intra- and inter-individual variation in the specific observations made by participants, which calls into question the reliability of dorsal hand traits as suitable points of interest for photographic hand comparison." Whatever Boyd et al. meant by "specific features" and "minimal ... disagreement," they somehow concluded that "[t]hese findings are consistent with recent studies that show support for qualitative methods of identification and have implications for current efforts to develop quantitative methods based off the traits investigated here."
  9. After the report was published, Propublica reporter Ryan Gabrielson kindly called my attention to the transcript of the direct examination on the fourth day of the trial and the cross-examination on the fifth day, which Propublica has made available as (low resolution) pdf files.
POSTINGS IN THIS SERIES
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 1), Mar. 3, 2019
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 2), Mar. 19, 2019.
  • Propublica's Picture of Photographic Analysis at the FBI Laboratory (pt. 3), Mar. 20, 2019