Showing posts with label ASTM. Show all posts
Showing posts with label ASTM. Show all posts

Friday, September 3, 2021

Does Qualitative Measurement Uncertainty Exist?

I have heard it said that forensic-science standards for interpreting the results of chemical or other tests need not discuss uncertainty in measurements of qualitative properties. For instance, ASTM International appropriately requires standards for test methods to include a section reporting on precision and bias as manifested in interlaboratory tests. Yet, it applies this requirement exclusively to quantitative measurements. Its 2021 style manual is unequivocal:

When a test method specifies that a test result is a nonnumerical report of success or failure or other categorization or classification based on criteria specified in the procedure, use a statement on precision and bias such as the following: “Precision and Bias—No information is presented about either the precision or bias of Test Method X0000 for measuring (insert here the name of the property) since the test result is nonquantitative" (ASTM 2020, § A21.5.4, pp. A3-A14).

Qualitative measurements are observation-statements such as the ink is blue, the friction ridge skin pattern includes loops, the bloodstain displays a cessation pattern, the blood group is type A, the glass fragments fit together perfectly, or the material contains cocaine. Likewise, the statements could be comparative: the recording of an unknown bell ringing sounds like it has a higher pitch than the ringing of a known bell; the hairs are microscopically indistinguishable; or the striations on the recovered bullet and the test bullet line up when viewed in the comparison microscope.

“Precision” is defined as “the closeness of agreement between test results obtained under prescribed conditions” (ibid. § A21.2.1, at A12). “A statement on precision allows potential users of the test method to assess in general terms its usefulness in proposed applications” and is mandatory (ibid. § A21.2, at A12). So how can it be that statements of precision and bias are not allowed for qualitative as opposed to quantitative findings? In both situations, the system that generates the findings could be noisy or skewed in its outcomes.

The only answer I have heard is that measurements cannot be qualitative because the word "measurement" is reserved for determining the magnitude of quantities such as length or mass. The values of these quantitative variables are basically isomorphic to the nonnegative real numbers. Counts, such as the number of alpha particles emitted in a given interval of time by radium atoms, also qualify as measurements because there is a quantitative, additive structure to them. The values of the variable are basically isomorphic to the natural numbers. Properties that only have names are described by nominal variables. Although numbers can assigned (1 for a match and 0 for a nonmatch, for example) these numbers are no more a measurement than a social security number is. In short, the argument is that because “measurements” do no not include qualitative judgments, classifications, decisions, identifications, or whatever one might call them, no statement of measurement uncertainty or error is possible, let alone required.

This argument is incredibly weak. To begin with, the definition of “measurement” is a highly contested concept. As one guide from NIST explains, a “much wider” conception of measurement than the one “contemplated in the current version of the International vocabulary of metrology (VIM)” has been developed in the metrology literature, and the measurand “may be ... qualitative (for example, the provenance of a glass fragment determined in a forensic investigation" (Possolo 2015). Broader conceptions of measurement have been the subject of many decades of writing in psychology and psychometrics (see, e.g., Humphry 2017; Mitchell 1990). Philosophers have been struggling to describe the scope and meaning of "measurement" at least since Aristotle (see, e.g., Tal 2015).

Second, even if one agrees with the definition in one NIST publication that “[m]easurement is [confined to] an experimental process that produces a value that can reasonably be attributed to a quantitative property of a phenomenon, body, or substance” (NIST 2019), some qualitative observations fit this definition. The color of a strip of litmus paper, for instance, can be understood as a value “that can reasonably be attributed to a quantitative property,” It is simply a crude measurement of pH.

Finally, the argument that there can be no measurement error for qualitative properties because those properties are not really “measured” is a semantic ploy that misses the point. The observations or estimates of nonquantitative properties as well as the individual measurements of quantitative properties are all subject to possible random and systematic error, and statements expressing the range of probable error for all measurements, observations, estimates, and classifications are essential. The need for these statements cannot be avoided for qualitative properties or judgments by the fiat of the VIM or some other dictionary. Even if “measurement” must be read in one particular, narrow, technical sense, “evaluation uncertainty” or “examination uncertainty” still must be reckoned with (Mari et al. 2020).

In sum, there is no excuse for ASTM and other organizations promulgating standards for forensic-science test methods to exempt any reported findings from required statements of uncertainty. Many statistics can be used to indicate how reliable (repeatable and reproducible) and valid (accurate) the test results may be (ibid.; Ellison & Gregory 1998; Pendrill & Petersson 2016). The qualitative-quantitative distinction affects the choice of the statistical method or expression but not the need to have one.

REFERENCES

  • ASTM Int’l, Form and Style for ASTM Standards (2020), https://www.astm.org/FormStyle_for_ASTM_STDS.html.
  • Stephen L. R. Ellison & Soumi Gregory, Perspective: Quantifying Uncertainty in Qualitative Analysis, Analyst 123, 1155-1161 (1998), https://doi.org/10.1039/A707970B
  • Stephen M. Humphry, Psychological Measurement: Theory, Paradoxes, and Prototypes, 27(3) Theory & Psychology 407–418 (2017)
  • L. Mari, C. Narduzzi, S. Trapmann, Foundations of Uncertainty in Evaluation of Nominal properties, 152 Measurement 107397 (2020), DOI:10.1016/j.measurement.2019.107397
  • Joel Mitchell, An Introduction to the Logic of Psychological Measurement (1990)
  • NIST, Statistical Engineering Division, Measurement Uncertainty, updated Nov. 15, 2019, https://www.nist.gov/itl/sed/topic-areas/measurement-uncertainty
  • Leslie Pendrill & Niclas Petersson, Metrology of human-based and other qualitative measurements, 27(9) Measurement Sci. Technol. 27 094003 (2016)
  • A. Possolo, Simple Guide for Evaluating and Expressing the Uncertainty of NIST
    Measurement Results (NIST Technical Note 1900), 2015, doi: 10.6028/NIST.TN.1900
  • Eran Tal, Measurement in Science, in Stanford Encyclopedia of Philosophy (Edward N. Zalta ed. 2015), https://plato.stanford.edu/archives/fall2017/entries/measurement-science/

APPENDIX: ADDITIONAL PUBLICATIONS ON "QUALITATIVE MEASUREMENT"

  1. Mary J. Allen & Wendy M. Yen, Introduction to Measurement Theory 2 (1979) ("In measurement, numbers are assigned systematically and can be of various forms. For example, labeling people with red hair "1" and people with brown hair "2" is a measurement. Since numbers are assigned to individuals in a systematic way and differences between scores represent differences in the property being measured (hair color).")
  2. Peter-Th. Wilrich, The determination of precision of qualitative measurement methods by interlaboratory experiments, Accreditation and quality assurance, 15: 439-444 (2010)
  3. Boris L. Milman, Identification of chemical compounds, Trends in Analytical Chemistry, 24:6, 2005 ("identification itself is considered as measurement on a qualitative scale")
  4. NIST Expert Working Group on Human Factors in Latent Print Analysis, Latent Print Examination and Human Factors: Improving the Practice Through a Systems Approach, Gaithersburg: National Institute of Standards and Technology, David H. Kaye ed., 2012 (defining "measurement" broadly, to encompass categorical variables, including the examiner's judgment about the source of a print).
  5. Lim, Yong Kwan, Kweon, Oh Joo, Lee, Mi-Kyung and Kim, Hye Ryoun. Assessing the measurement uncertainty of qualitative analysis in the clinical laboratory. Journal of Laboratory Medicine, vol. 44, no. 1, 2020, pp. 3-10. https://doi.org/10.1515/labmed-2019-0155 ("Measurement uncertainty is a parameter that is associated with the dispersion of measurements. Assessment of the measurement uncertainty is recommended in qualitative analyses in clinical laboratories; however, the measurement uncertainty of qualitative tests has been neglected despite the introduction of many adequate methods.")
  6. Donald Richards, Simultaneous Quantitative and Qualitative Measurements in Drug-Metabolism Investigations, Pharmaceutical Technology 2013
  7. Kadri Orro, Olga Smirnova, Jelena Arshavskaja, Kristiina Salk, Anne Meikas, Susan Pihelgas, Reet Rumvolt, Külli Kingo, Aram Kazarjan, Toomas Neuman & Pieter Spee, Development of TAP, a non-invasive test for -qualitative and quantitative measurements of biomarkers from the skin surface, Biomarker Research 2: 20 (2014)
  8. J M Conly & K Stein, Quantitative and qualitative measurements of K vitamins in human intestinal contents, Am J Gastroenterol. 1992 Mar;87(3):311-316
  9. Wenjia Meng, Qian Zheng, Gang Pan, Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network, IEEE Transactions on Neural Networks and Learning Systems 2020
  10. Rudolf M. Verdaasdonk, Jovanie Razafindrakoto, Philip Green, Real time large scale air flow imaging for qualitative measurements in view of infection control in the OR (Conference Presentation) Proceedings Volume 10870, Design and Quality for Biomedical Technologies XII; 1087002 (2019) https://doi.org/10.1117/12.2511185
  11. Rashis, Bernard, Witte, William G. & Hopko, Russell N., Qualitative Measurements of the Effective Heats of Ablation of Several Materials in Supersonic Air Jets at Stagnation Temperatures Up to 11,000 Degrees F, National Advisory Committee for Aeronautics, July 7, 1958
  12. Lawrence F Cunningham and Clifford E Young, Quantitative and Qualitative Approaches, Journal of Public Transportation 1(4) (1997) ("The study also contrasts the results of quantitative and qualitative measurements and methodologies for assessing transportation service quality")
  13. JM Conly, K Stein, Quantitative and qualitative measurements of K vitamins in human intestinal contents, American Journal of Gastroenterology, 1992
  14. P Sinha, Workshop on Biologically Motivated Computer Vision, 2002 - Springer ("Our emphasis on the use of qualitative measurements renders the representations stable in the presence of sensor noise and significant changes in object appearance. We develop our ideas in the context of the task of face-detection under varying illumination")
  15. D Michalski, S Liebig, E Thomae & A Hinz, Pain in Patients with Multiple Sclerosis: a Complex Assessment Including Quantitative and Qualitative Measurements, 40 J. Pain 219–225 (2011), https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3160835/
  16. Cécilia Merlen, Marie Verriele, Sabine Crunaire,Vincent Ricard, Pascal Kaluzny, Nadine Locoge, Quantitative or Only Qualitative Measurements of Sulfur Compounds in Ambient Air at Ppb Level? Uncertainties Assessment for Active Sampling with Tenax TA®, 132 Microchemical J. 143-153 (2017)
  17. Tomomichi Suzuki, Jun Ichi Takeshita, Mayu Ogawa, Xiao-Nan Lu, Yoshikazu Ojima, Analysis of Measurement Precision Experiment with Categorical Variables, 13th International Workshop on Intelligent Statistical Quality Control 2019, Hong Kong ("Evaluating performance of a measurement method is essential in metrology. Concepts of repeatability and reproducibility are introduced in ISO5725-1 (1994) including how to run and analyse experiments (usually collaborative studies) to obtain these precision measures. ISO5725-2 (1994) describe precision evaluation in quantitative measurements but not in qualitative measurements. Some methods have been proposed for qualitative measurements cases such as Wilrich (2010), de Mast & van Wieringen (2010), Bashkansky, Gadrich & Kuselman (2012). Item response theory (Muraki, 1992) is another methodology that can be used to analyse qualitative data.").

Saturday, March 6, 2021

"The Judgment of an Experienced Examiner"

The quotation from a textbook on forensic science in the left-hand panel invites the question of what the author was thinking. That an examiner's judgment is more important in comparisons of "class" features than "individual" ones? That the features that are present are less important than the examiner's judgment of them? Neither interpretation makes much sense. The examiner's judgment does not dictate anything about the features. It is the other way around. In applying a valid method, a proficient examiner generally will make correct judgments of significance as dictated by the features that are present.

As with most class evidence, the significance of a fiber comparison is dictated by the circumstances of the case, by the location, number, and nature of the fibers examined, and, most important, by the judgment of an experienced examiner.
Richard Saferstein, Criminalistics: An Introduction to Forensic Science 272 (12rth ed. 2018, Pearson Education Inc.) (emphasis added)
Over the years, scientific and legal scholars have called for the implementation of algorithms (e.g., statistical methods) in forensic science to provide an empirical foundation to experts’ subjective conclusions. ... Reactions have ranged from passive skepticism to outright opposition, often in favor of traditional experience and expertise as a sufficient basis for conclusions. In this paper, we explore why practitioners are generally in opposition to algorithmic interventions and how their concerns might be overcome. We accomplish this by considering issues concerning human-algorithm interactions in both real world domains and laboratory studies as well as issues concerning the litigation of algorithms in the American legal system. [W]e propose a strategy for approaching the implementation of algorithms ... .
Henry Swofford & Christophe Champod, Implementation of Algorithms in Pattern & Impression Evidence: A Responsible and Practical Roadmap, 3 Forensic Sci. Int'l: Synergy 100142 (2021) (abstract)

Another example of sloppy phrasing about expert judgment is the boilerplate disclaimers or admonitions in ASTM standards for forensic-science methods. For example, the 2019 Standard Guide for Forensic Analysis of Fibers by Infrared Spectroscopy (E2224−19) insists (in italics no less) that

This standard cannot replace knowledge, skills, or abilities acquired through education, training, and experience and is to be used in conjunction with professional judgment by individuals with such discipline-specific knowledge, skills, and abilities.

Does the first independent clause mean that fiber analysts are free to depart from the standard on the basis of their general "knowledge, skills, or abilities acquired through education, training, and experience"? Ever since Congress funded the Organization of Scientific Area Committees for Forensic Science (OSAC) to write new and better standards, lawyers in the organization have objected to the ASTM wording (without doubting that expert methods should be applied by responsible experts).

Now rumor has it that ASTM will be changing its stock sentence to the less ambiguous observation that

This standard is intended for use by competent forensic science practitioners with the requisite formal education, discipline-specific training (see Practice E2917), and demonstrated proficiency to perform forensic casework.

That's innocuous. Indeed, you might think it goes without saying.

Saturday, May 6, 2017

Who Copy Edits ASTM Standards?

This posting is not about science or law. It is about English writing. I recently had occasion to read the “Standard Guide for Analysis of Clandestine Drug Laboratory Evidence” issued by ASTM International, a private standards development organization. The standard exemplifies a common problem with the ASTM standards for forensic science — an apparent absence of copy and line editing to achieve clear and efficient expression of the ideas of the committees that write the standards. 1/

This particular standard, known as E2882-12, opens with an observation about the “scope” of the document — namely, that
This guide does not replace knowledge, skill, ability, experience, education, or training and should be used in conjunction with professional judgment.
The word “replace” has caused a couple of readers to complain that this admonition implies that unstructured “knowledge, skill, ability, experience, education, or training” suffices for the analysis of the evidence. That is not  a fair reading of the sentence, but joining the two clauses with “and” makes it seem like they are separate points. Why not make it as easy as possible for the reader to get the intended message? I think the sentence amounts to nothing more than the following simple idea:
This standard is intended to help professionals use their knowledge and skill to analyze clandestine drug laboratory evidence.
Why not just say this? Why all the extra verbiage?

Unfortunately, this text is not an isolated example of the need for detailed editing. Another infelicity is
capacity—the amount of finished product that could be produced, either in one batch or over a defined period of time, and given a set list of variables.
The words “and given a set list of variables” are a sentence fragment. They dangle aimlessly after the comma. The copy edit is obvious:
Capacity is the amount of finished product that could be produced, for a specified set of variables, either in one batch or over a stated period of time.
It still may not be clear what a “set of variables” means here, but at least the words about unnamed variables occur where they belong.

The wording in a section on reporting is especially obscure:
Laboratories should have documented policies establishing protocols for reviewing verbal information and conclusions should be subject to technical review whenever possible. It is acknowledged that responding to queries in court or investigative needs may present an exception.
One clear statement of what the sentences seem to assert is that
Laboratories should have written protocols to ensure that oral communications from laboratory personnel are reviewed for technical correctness. However, a protocol can dispense with (1) review of some courtroom testimony and (2) review that would impede an investigation.
Whether this edited version expresses what the authors wanted to say or presents a satisfactory policy is unclear, but at least the version is more easily understood.

Other phrases that should raise red flags for editing abound. I’ll end with three examples.
  • This guide does not purport to address all of the safety concerns, if any, associated with its use. The editor would say: Make up your mind. If there are no safety concerns, then the sentence is worthless. If there are safety concerns, then the standard should address them. If there is a reason not to address all of them, then the standard can say, “There are additional safety concerns for a laboratory to consider.” If there is a desire to be very cautious, it could read, “There could be additional safety concerns for a laboratory to consider.”
  • ... calculations can be achieved from ... . Copy editor: It sounds odd to speak of "achieving" calculations. The phrase "calculations can be made by" would be more apt.
  • Quantitative measurements of clandestine laboratory samples have an accuracy which is dependent on sampling and, if a liquid, on volume calculations. This sentence is both circumlocutious ("which is dependent") and disjointed ("if a liquid" is in the wrong place to modify "samples"). It also seems to conflate measurements on subsamples of the material submitted for analysis ("clandestine laboratory samples") with inference from the subsamples to the sample of the seized items. If this reading of the dense sentence is correct, editing would expand it along the following lines: "The accuracy of quantitative measurements of a liquid sample depends on the calculated volume of the sample. When the material analyzed is not the entire sample, then the accuracy of any inferences to the entire sample also depends on the homogeneity of the sample and the procedure by which the subsample was chosen.
Good writing requires the right words in the correct order. Good editing makes the writing more readable. Many existing technical standards in forensic science still need good editing to make them fully fit for purpose.

NOTE
  1. Although some publishers distinguish between line editing and copy editing, this posting uses the phrase "copy editing" broadly, to refer to the process of reviewing and correcting written material to ensure "that whatever appears in public is accurate, easy to follow, and fit for purpose." Society for Editors and Proofreaders, FAQs: What Is Copy-editing?, https://www.sfep.org.uk/about/faqs/what-is-copy-editing/.

Saturday, April 2, 2016

Sample Evidence: What’s Wrong with ASTM E2548-11 Standard Guide for Sampling Seized Drugs?

Samuel Johnson once observed that “You don't have to eat the whole ox to know that it is tough.” Or maybe he didn't say this, 1/ but the idea applies to many endeavors. One of them is testing of seized drugs. The law needs to—and generally does—recognize the value of surveys and samples in drug and many other kinds of cases. 2/ If the quantity of seized drugs is large, it is impractical and typically unnecessary to test every bit of the materials. Clear guidance on how to select samples from the population of seized matter would be helpful to courts and laboratories alike.

To accomplish this goal, the Chemistry and Instrumental Analysis Subject Area Committee of the Organization of Scientific Area Committees for Forensic Science (OSAC) has recommended the addition of ASTM International’s Standard Guide for Sampling Seized Drugs for Qualitative and Quantitative Analysis (known as ASTM E2548-11) to the National Institute of Standards and Technology (NIST) Registry of Approved Standards. Unfortunately, this "Standard Guide" is vague in its guidance, incomplete and out of date in its references, and nonstandard in its nomenclature for sampling.

The Standard does not purport to prescribe "specific sampling strategies." 3/ Instead, it instructs “[t]he laboratory ... to develop its own strategies” and “recommend[s] that ... key points be addressed.” There are only two key points. 4/ One is that “[s]tatistically selected units shall be analyzed to meet Practice E2329 if statistical inferences are to be made about the whole population.” 5/ But ASTM E2329 merely describes the kinds of analytical tests that can or should be performed on samples. It reveals nothing about how to draw samples from a population. So far, ASTM E2548 offers no guidance about sampling.

The other “key point” is that “[s]ampling may be statistical or non-statistical.” Although tautological (A is either X or not-X), X is not defined, and an explanatory note intensifies the ambiguity. It states that “[f]or the purpose of this guide, the use of the term statistical is meant to include the notion of an approach that is probability-based.” 6/  Does “probability-based” mean probability sampling (the subject of ASTM E105-10)? At least the latter has a well-defined meaning in sampling theory. 7/ It means that every unit in the sampling frame has a known probability of being drawn.

But even if this is what ASTM E2548-11 Standard Guide means by “probability-based,” the phrase is not congruent with "statistical." The note indicates that even sampling that is not “probability-based” still can be considered "statistical sampling." Later parts of the the Standard allow inferences to populations to be made from "statistical" samples but not from "non-statistical" ones. Using an undefined notion of "statistical" and "non-statistical" as the fundamental organizing principle departs from conventional statistical terminology and reasoning. The usual understanding of sampling differentiates between probability samples -- for which sampling error readily can be quantified -- and other forms of sampling (whether systematic or ad hoc) -- for which statistical analysis depends on the assumption that the sample is the equivalent of a probability sample.

Thus, the statistical literature on sampling commonly explains that
If the probability of selection for each unit is unknown, or cannot be calculated, the sample is called a non-probability sample. Non-probability samples are often less expensive, easier to run and don't require a frame. [¶] However, it is not possible to accurately evaluate the precision (i.e., closeness of estimates under repeated sampling of the same size) of estimates from non-probability samples since there is no control over the representativeness of the sample. 8/
In contrast, because the ASTM Standard does not focus on probability sampling as opposed to other "statistical sampling," the laboratory personnel (or the lawyer) reading the standard never learns that "it is dangerous to make inferences about the target population on the basis of a non-probability sample." 9/

Indeed, Figure 1 of ASTM 2548 introduces further confusion about "statistical sampling." In this figure, a statistical “sampling plan” is either “Hypergeometric,” “Bayesian,” or “Other probability-based.” But the sampling distribution of a statistic is not a “sampling plan” (although it could inform one). A sampling plan should specify the sample size (or a procedure for stopping the sampling if results on the sampled items up to that point make further testing unnecessary). For sampling from a finite population without replacement, the hypergeometric probability distribution applies to sample-size computations and estimates of sampling error. But how does that make the sampling plan hypergeometric? One type of “sampling plan” would be to draw a simple random sample of a size computed to have a good chance of producing a representative sample. Describing a plan for simple random sampling, stratified random sampling, or any other design as “hypergeometric,” “Bayesian,” or “other” is not helpful.

Similarly confusing is the figure’s trichotomy of “non-statistical” into the following “plans”: “Square root N,” “Management Directive,” and “Judicial Requirements.” Using the old √N + 1 rule of thumb for determining sample size may be sub-optimal, 10/ but it is “statistical” -- it uses a statistical computation to establish a sample size. So do any judicial or administrative demands to sample a fixed percentage of the population (an approach that a Standard should deprecate). No matter how one determines the sample size, if probability sampling has been conducted, statistical inferences and estimates have the same meaning.

Also puzzling are the assertions that “[a] population can consist of a single unit,” 11/ and that “numerous sampling plans ... are applicable to single and multiple unit populations.” 12/ If a population consists of “a single unit” (as the term is normally used), 13/ then a laboratory that tests this unit has conducted a census. The study design does not involve sampling, so there can be no sampling error.

When it comes to the issue of reporting quantities such as sampling error, the ASTM Standard is woefully inadequate. The entirety of the discussion is this:
7.1 Inferences based on use of a sampling plan and concomitant analysis shall be documented.

8.1 Sampling information shall be included in reports.
8.1.1 Statistically Selected Sample(s)—Reporting statistical inferences for a population is acceptable when testing is performed on the statistically selected units as stated in 6.1 above [that is, according to a standard that is on the NIST Registry with a disclaimer by NIST]. The language in the report must make it clear to the reader that the results are based on a sampling plan.
8.1.2 Non-Statistically Selected Sample(s)—The language in the report must make it clear to the reader that the results apply to only the tested units. For example, 2 of 100 bags were analyzed and found to contain Cocaine.
These remarks are internally problematic. For example, why would an analyst report the population size, the sample size, and the sample data for “non-statistical” samples but not for “statistical” ones?

More fundamentally, to be helpful to the forensic-science and legal communities, a standard has to consider how the results of the analyses should be presented in a report and in court. Should not the full sampling plan be stated — the mechanism for drawing samples (e.g., blinded, which the ASTM Standard calls “black box” sampling, or selecting from numbered samples by a table of random numbers, which it portrays as not “practical in all cases”); the sample size; and the kind of sampling (simple random, stratified, etc.)? It is not enough merely to state that “the results are based on a sampling plan.”

When probability sampling has been employed, a sound foundation for inferences about population parameters will exist. But how should such inference be undertaken and presented? A Neyman-Pearson confidence interval? With what confidence coefficient? A frequentist test of a hypothesis? Explained how? A Bayesian conclusion such as “There is a probability of 90% that the weight of the cocaine in the shipment seized exceeds X”? The ASTM Standard seems to contemplate statements about “[t]he probability that a given percentage of the population contains the drug of interest or is positive for a given characteristic,” but it does not even mention what goes into computing a Bayesian credible interval or the like. 14/

The OSAC Newsletter proudly states that "[a] standard or guideline that is posted on either Registry demonstrates that the methods it contains have been assessed to be valid by forensic practitioners, academic researchers, measurement scientists, and statisticians through a consensus development process that allows participation and comment from all relevant stakeholders." The experience with ASTM Standards 2548 and 2329 suggests that even before a proposed standard can be approved by a Scientific Area Committee, the OSAC process should provide for a written review of statistical content by a group of statisticians. 15/

Disclosure and disclaimer: I am a member of the OSAC Legal Resource Committee. The information and views presented here do not represent those of, and are not necessarily shared by NIST, OSAC, any unit within these organizations, or any other organization or individuals.

Notes
  1. According to the anonymous webpage Apocrypha: The Samuel Johnson Sound Bite Page, the aphorism is "apocryphal because it's not found in his works, letters, or contemporary biographies about Samuel Johnson. But it is similar to something he once said about Mrs. Montague's book on Shakespeare: 'I have indeed, not read it all. But when I take up the end of a web, and find it packthread, I do not expect, by looking further, to find embroidery.'"
  2. See, e.g., David H. Kaye et al., David E. Bernstein & Jennifer L. Mnookin, The New Wigmore: A Treatise on Evidence: Expert Evidence (2d ed. 2011); Hans Zeisel & David H. Kaye, Prove It with Figures: Empirical Methods in Law and Litigation (1997).
  3. See ASTM E2548-11, § 4.1. 
  4. Id., § 4.2.
  5. § 4.2.2.
  6. § 4.2.1 (emphasis added).
  7. E.g., Statistics Canada, Probability Sampling, July 23, 2013:
    Probability sampling involves the selection of a sample from a population, based on the principle of randomization or chance. Probability sampling is more complex, more time-consuming and usually more costly than non-probability sampling. However, because units from the population are randomly selected and each unit's probability of inclusion can be calculated, reliable estimates can be produced along with estimates of the sampling error, and inferences can be made about the population.
  8. National Statistical Service (Australia), Basic Survey Design, http://www.nss.gov.au/nss/home.nsf/SurveyDesignDoc/B0D9A40C6B27487BCA2571AB002479FE?OpenDocument (emphasis in original).
  9. Id.
  10. See J. Muralimanohar & K. Jaianan, Determination of Effectiveness of the “Square Root of N Plus One” Rule in Lot Acceptance Sampling Using an Operating Characteristic Curve, Quality Assurance Journal, 14(1-2): 33.37, 2011.
  11. § 5.2.2.
  12. § 5.3.
  13. Laboratory and Scientific Section, United Nations Office on Drugs and Crime, Guidelines on Representative Drug Sampling 3 (2009).
  14. Cf. James M. Curran, An Introduction to Bayesian Credible Intervals for Sampling Error in DNA Profiles, Law, Probability and Risk, 4, 115−126, 2011, doi:10.1093/lpr/mgi009
  15. Of course, no process is perfect, but early statistical review can make technical problems more apparent. Cf. Sam Kean, Whistleblower Lawsuit Puts Spotlight On FDA Technical Reviews, Science, Feb. 2, 2012.

Saturday, March 26, 2016

NIST Distances Itself from the First OSAC-approved Forensic Science Standard

On January 11, 2016, a group created by the federal government to develop better standards for forensic science approved -- without changing a single word of substance -- a standard previously promulgated by a committee of ASTM International (formerly the American Society for Testing Materials). The federally mandated body that showcased this standard is the Organization of Scientific Area Committees for Forensic Science (OSAC). It is supported, administratively and financially, by the Department of Commerce's highly respected National Institute of Standards and Technology (NIST). 1/

The approved standard has the ponderous name of ASTM E2329−14 Standard Practice for Identification of Seized Drugs. To my eye, it looks like an odd choice for the first (and thus far, only) entry in the OSAC Registry of Approved Standards. Why so?

For one thing, this one standard itself seems to approve of nine other ASTM standards -- none of which have been vetted by OSAC. Second, the FSSB approved the standard over objections from two out of the three OSAC "resource committees" -- its Legal Resources Committee and its Human Factors Committee. 2/ Third. the standard permits definitive conclusions based on standardless, subjective assessments of botanical specimens. Fourth, without discussing or citing any studies of error probabilities for the methods involved, the standard states or suggests that false-positive errors will not occur. Thus, an earlier posting tartly contrasted some of the language in the Standard to the admonition in the 2009 report of the National Research Council Committee on Identifying the Needs of the Forensic Science Community that "[a]ll results for every forensic science method should indicate the uncertainty in the measurements that are made, and studies must be conducted that enable the estimation of those values."

On March 17, 2016, more than two months after the NIST-created OSAC adopted this first standard, NIST issued a public statement disavowing the standard as written because "concerns have been raised that some of the language in the standard is not scientifically rigorous." 3/ Like the National Research Council, NIST appreciates that "no measurement, qualitative or quantitative, should be characterized as without the risk of error or uncertainty." 4/ The statement adds that "NIST and the FSSB have independently asked that ASTM review the language."

The FSSB's action is, in a way, quite puzzling. Why would the FSSB want the organization that already wrote and approved the standard that the FSSB reviewed and adopted as a registry-ready, gold standard to "review" the FSSB-approved standard? It is not as if some new scientific research suddenly undermined the standard, requiring it to be revised. And even if new information had surfaced after the FSSB voted, the appropriate response would have been to take it down until the issue could be resolved.

Moreover, why would NIST and the FSSB "independently" ask for ASTM review when the subcommittee that wanted the standard on the registry already had promised to secure revisions through ASTM. The record before the FSSB included the following response to criticisms filed by the Legal Resource Committee:
The Seized Drug subcommittee intends to clarify the quoted language pertaining to uncertainty and error during the next ASTM revision of this document." E2329-14 Seized Drugs Response to LRC Comments FINAL.pdf (277K) SAC Chemistry/Instrument Analysis, Jan. 11, 2016
Does the fact that NIST and the FSSB have added their voices to that of the OSAC Subcommittee on Seized Drugs mean that revisions that clearly should have been made before posting a standard to the repository will occur any sooner? And what can justify leaving a standard that admittedly needs "clarification" on the registry pending the requisite rewriting? Cannot laboratories continue to use the ASTM standard to guide them just as they did before the OSAC registry existed?

Whatever the answers to these questions may be, NIST's reservations about the first OSAC standard, although not spelled out in full, were the subject of questions at a recent meeting of the National Commission on Forensic Science. On March 21, 2016, Commissioner Marc LeBeau asked presenters from NIST whether NIST planned to post statements of agreement as well as disagreement for every future OSAC-approved standard. I cannot locate a transcript or videotape of the meeting, but my recollection is that the answer was essentially "no."

No doubt, NIST hopes that the kerfuffle over ASTM E2329−14 is one off, but the apparent inclination of OSAC subcommittees to try to import unimproved ASTM standards into the registry does not bode well. The latest example is ASTM E2388-11 Standard Guide for Minimum Training Requirements for Forensic Document Examiners. It is up for consideration as an OSAC standard (and for public comment during the next couple of weeks) even though OSAC has no approved standard on what the document examiners are expected to do once they are trained.

Disclosure and disclaimer: I am a member of the OSAC Legal Resource Committee. The information and views presented here do not represent those of, and are not necessarily shared by, NIST, OSAC, any unit within these organizations, or any other organization or individuals.

Notes
  1. See OSAC Subcommittee and Scientific Area Committee (SAC) Chairs, https://rticqpub1.connectsolutions.com/content/connect/c1/7/en/events/event/shared/1187757659/speaker_info.html?sco-id=1187765255 ("OSAC is part of an initiative by NIST and the Department of Justice to strengthen forensic science in the United States. The organization is a collaborative body of more than 500 forensic science practitioners and other experts who represent local, state, and federal agencies; academia; and industry. NIST has established OSAC to support the development and promulgation of forensic science consensus documentary standards and guidelines, and to ensure that a sufficient scientific basis exists for each discipline.").
  2. The third resource committee, the Quality Infrastructure Committee (QIC), does not seem to comment on the substance of proposed standards.
  3. NIST Statement on ASTM Standard E2329-14, Mar. 17, 2016, http://www.nist.gov/forensics/nist-statement-on-astm-e2329-14.cfm
  4. That said, the NIST statement cautions that "It is important to note that NIST is not contesting results obtained from seized evidence using the standard." Id.

Monday, March 7, 2016

Hot Paint: Another ASTM Standard (E2937-13) that Needs More Work

A second standard for comparing samples of paint under review for the OSAC Registry is ASTM E2937-13, on "Infrared Spectroscopy in Forensic Paint Examinations." It raises several of the issues previously noted in the broader ASTM E1610-14 "Standard Guide for Forensic Paint Analysis and Comparison."

The standard for IR spectroscopy presupposes that the goal is to “to determine whether any significant differences exist between the known and questioned samples,” where a “significant difference” is “a difference between two samples that indicates that the two samples do not have a common origin.” The criminalist then is expected to declare whether “[s]pectra are dissimilar,” “indistinguishable,” or “inconclusive.”

Although categorical judgments have the benefit of simplicity and familiarity, most literature on forensic inference now maintains that analysts should present statements about the weight of the evidence rather than categorical conclusions about source hypotheses. By considering and presenting the degree to which the observations support one hypothesis as compared to another without dictating the conclusion that must be drawn, the analyst supplies the most information. It is not clear whether the standard rejects this view and is intended to preclude experts from using a weight-of-evidence approach to the comparison process.

The categorical approach that the standard adopts is notable for its vagueness. On its face, the definition of “significant difference” permits analysts to declare that differences with almost no discriminating power are so significant that two samples “do not have a common origin.” This lack of guidance arises because any difference that occurs more frequently among two samples with different origins than among two same-source samples “indicates” different origins and hence is “significant.” For example, a difference that arises 1,000 times more often for different-source samples is indicative of difference sources. But so is a difference that arises only 10% more often for different-source samples. Both “indicate” non-association. They differ only in the magnitude of the measure of non-association. The 1,000-times-more-often quantity is a strong indication of non-association, whereas the 10% figure is a weak indication. But in both cases, the differences indicate (to some degree) non-association relative to association.

To avoid this looseness, one might try to read “indicates” as connoting “strongly indicates” or “establishes,” but there is no reason to promulgate an ambiguous standard that requires readers in the fields of forensic science and law to struggle to discern and supply its intended meaning. And, if “establishes” is the intended meaning, then more guidance is needed to help analysts determine, on the basis of objective data about the range of differences seen in same-source and in different-source samples, when a difference is “significant” in the sense of discriminating between the former and the latter types of samples. That is, the standard should supply a validated decision rule; it should present the conditional error probabilities of this decision rule; and it should refer specifically to the studies that have validated it. These features of standards are not absolute requirements for admitting scientific evidence, but they would go far to assuring courts and counsel that the criteria of “known or potential rate of error” and “standards controlling the technique's operation” enumerated in Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579, 594 (1993), militate in favor of admissibility (and persuasiveness of the testimony if a case goes to trial).

Section 10.6.1.1 of the ASTM Standard does not begin to do this. It offers an unbounded “rule of thumb” — “that the positions of corresponding peaks in two or more spectra be within ±5 cm^-1. For sharp absorption peaks one should use tighter constraints. One should critically scrutinize the spectra being compared if corresponding peaks vary by more than 5 cm^-1. Replicate collected spectra may be necessary to determine reproducibility of absorption position.” What is the basis for “critical scrutiny”? How many replicates are necessary? When are they necessary? What is the accuracy of examiners who follow the open-ended “rule of thumb”?

Given the lack of standards for deciding what is “significant,” the definitions of “dissimilar,” “indistinguishable,” and “inconclusive” are indeterminate. They read:
  • 10.7.1 Spectra are dissimilar if they contain one or more significant differences.
  • 10.7.2 Spectra are indistinguishable if they contain no significant differences.
  • 10.7.3 A spectral comparison is inconclusive if sample size or condition precludes a decision as to whether differences are significant.
Inasmuch as any difference can be considered “significant,” the criminalist has no basis in the standard to declare an inclusion, an exclusion, or an inconclusive outcome. This deprives the standard of the legally desirable status under Daubert of “standards controlling the technique's operation.”

Wednesday, February 24, 2016

Is OSAC Painting Itself Out of the Picture? Time to Comment on ASTM E1640-14

The Organization of Scientific Area Committees (OSAC) started on the wrong foot with the first standard it chose to place on its blue-ribbon Registry of Approved Standards. Two more standards from the Chemistry and Instrumental Analysis Committee are currently up for public comment. Here, I discuss the first of the two, known in the field as ASTM E1610-14.

ASTM E1610-14 is a "Standard Guide for Forensic Paint Analysis and Comparison." It seeks "to assist individuals who conduct forensic paint analyses in their evaluation, selection, and application of tests that may be of value to their investigations." To a large extent, it accomplishes this goal.

At the same time, however, this Standard Guide fails to provide much useful guidance on a matter of critical concern to the legal system — reporting results in a way that fairly conveys their probative value and the inevitable uncertainty of any scientifically validated test. In fact, the Guide fails to reflect or acknowledge modern thinking about the interpretation of forensic-science test results.

The premise of this document seems to be that the main task of the criminalist is to make "physical matches between known and questioned samples" (sec. 7.1) based on "significance assessments" (sec. 8.3) in which a "significant difference" is "a difference between two samples that indicates that the two samples do not have a common origin." (Sec. 3.2.10). Although this definition of “significant” offers little or no guidance, these matches are expected to be "conclusive." (Sec. 8.6.1). This categorical approach does not represent the modern view of the evaluative statements that criminalists should make. Most literature on forensic inference now maintains that analysts should present statements about the weight of the evidence rather than categorical conclusions. 1/

If analysts studying traces of paint are to make "conclusive" statements, however, and if the NIST-supported OSAC organization is to follow through on the National Academy of Sciences' recommendation to incorporate estimates of uncertainty into forensic science, the conditional error probabilities for these conclusions must be provided.

The concept of validity in measurement is closely connected to estimating uncertainty. Section 1.2 announces that “[t]he need for validated methods and quality assurance guidelines is also addressed,” but this sentence is the only place in the document where the words "validity," "validated," or "validation" appear, and classifications based on a process with unknown sensitivity and specificity cannot be considered validated. Indeed, this ASTM Standard Guideline indirectly endorses a minority notion (from the legal perspective) on what it takes to establish validity. It approvingly refers (in sec. 4.1) to quality assurance guidelines from a 1999 SWGMAT document (actually published in 2000). These guidelines include the observation that "[t]echniques and procedures ... currently accepted by the scientific community should be considered valid." But it is widely appreciated that the general-acceptance criterion for scientific validity is not necessarily sufficient under Federal Rule of Evidence 702 and the rules of many states.

ASTM 1610-14 also contemplates testimony about a "physical match" (sec. 8.6) based on postulated "individualizing characteristics" rather than more appropriate probabilistic assessments. Section 8.6.1 advises that "[t]he most conclusive type of examination that can be performed on paint samples is physical matching. ... The corresponding features must possess individualizing characteristics."

The requirement of "individualizing characteristics" is too strict. Although intuition indicates that a combination of enough characteristics can constitute very strong evidence of a common source, it is not clear that any single characteristic is "individualizing." And, even if the existence of strictly individualizing characteristics has been demonstrated in the scientific literature, a criminalist should be able to use other characteristics in forming an expert opinion. After all, a combination of other characteristics that are known to be less than perfectly discriminating when considered one at a time can be highly discriminating when evaluated in toto.

Thus, the approach to "physical matches" in section 8.6.1 seems inconsistent with the logic of section 5.2 of the same ASTM document. This section explains that
Searching for differences between questioned and known samples is the basic thrust of forensic paint analysis and comparison. However, differences in appearance, layer sequence, size, shape, thickness, or some other physical or chemical feature can exist even in samples that are known to be from the same source. A forensic paint examiner’s goal is to assess the significance of any observed differences. The absence of significant differences at the conclusion of an analysis suggests that the paint samples could have a common origin. The strength of such an interpretation is a function of the type or number of corresponding features, or both.
This language is important for legal purposes because it affects the presentation of negative findings (“absence of significant differences”) that would incriminate a suspect or defendant. Surprisingly, there is nothing in the standard to guide or inform the analyst about how to report these results. Is the analyst expected to report only that “I could find no significant differences”? That seems insufficient. That “I am suggesting that the paint samples could have a common origin”? That even though “differences can exist even in samples that are known to be from the same source,” in this case there is “an absence of significant differences”? That may connote more than is appropriate. That, given “the type or number of corresponding features,” the interpretation of “common source” is very strong? What studies establish that a criminalist accurately can quantify (even verbally) the strength of the common-source interpretation? What is the uncertainty associated with such judgments? To be sure, the Standard Guide lists 69 references, but it is not clear which, if any, of them answer these questions.

Perhaps I am being too critical. Readers are invited to examine the ASTM Standard Guide for themselves and tell OSAC what they think.

Note
  1. E.g., John S. Buckleton et al., An Extended Likelihood Ratio Framework for Interpreting Evidence, 46 Sci. & Just. 69, 70 (2006) ("The idea of assessing the weight of evidence using a relative measure (known as the likelihood ratio) ... dominates the literature as the method of choice for interpreting forensic evidence across evidence types."); Angel Carrecedo, Forensic Genetics: History, in Forensic Biology 19, 22 (Max M. Houck ed. 2015) (“the single most important advance in forensic genetic thinking is the realization that the scientist should address the probability of the evidence.”).

Monday, February 15, 2016

Approximating Individualization: The ASTM's Standard Terminology for Digital Evidence

Forensic scientists portraying existing standards for evidence testing and evaluation have been known to praise the “rigorous standard development process of ASTM,” 1/ an internationally recognized standards development organization. Having looked over the organization's "Standard Terminology for Digital and Multimedia Evidence Examination" 2/ (not to mention some of its other standards), 3/ I wonder if the results are as rigorous (or as comprehensible) as they should be.

Consider the definition of "individualization":
Individualization, n—theoretically, a determination that two samples derive from the same source; practically, a determination that two samples derive from sources that cannot be distinguished within the sensitivity of the comparison process. (Compare identification.) DISCUSSION—Theoretical individualization is the asymptotic upper bound of the sensitivity of a source identification process.
The definition presents individualization as a theoretical construct that cannot be fully attained. But lots of things can be sorted down to the individual level—telephone, passport, and social security numbers are obvious examples. Surnames and given names are not individualizing at the national level, but they are among the students in almost every class that I have taught. As these examples suggest, individualization can only be defined for the elements of a set. 4/ If the set is enumerated and all its elements available for inspection, then it is possible to "individualize"—not just theoretically or "asymptotically," but practically and precisely.

Thus, the domain of the ASTM definition must be cases in which no exhaustive list of the elements is available. Even with this modification, however, the definition of "individualization" as "a determination that two samples derive from sources that cannot be distinguished within the sensitivity of the comparison process" is flawed for two reasons.

First, "sensitivity," not "specificity," must be what is intended. Sensitivity is the probability that the process will declare that an item comes from a source when it really does come from the source. The least upper bound on sensitivity (or any other probability) is 1. A process that always declares a positive association will have a sensitivity of 1 because it always will declare a positive association when there is one. The degree of source discrimination within a set of potential sources is the specificity. Only when the specificity equals 1 is exact individualization possible.

Second, the ASTM definition of "individualization" fails to state a crucial presupposition. If the specificity of the test verges on 1, then (by definition) the probability that a claim of individualization will be correct when the target item is the individual so identified also verges on 1. This is the approximate individualization that the ASTM is trying to define. But the definition as written does not require that the specificity be close to 1. An analyst following the words of the definition could claim to have "individualized" even when neither ideal nor approximate individualization exists.  As long as the specificity is not 1, "a determination that two samples derive from sources that cannot be distinguished" only shows that the item is an element of a class of indistinguishable items.

Notes
  1. Jay Siegel, Forensic Chemistry: Fundamentals and Applications 230 (2015).
  2. ASTM E2916-13, Standard Terminology for Digital and Multimedia Evidence Examination (2013), available for $44 at http://www.astm.org/Standards/E2916.htm.
  3. E.g., Broken Glass, Mangled Statistics, Forensic Science, Statistics & the Law, Feb. 3, 2016.
  4. David H. Kaye, Identification, Individuality, and Uniqueness: What's the Difference?, 8 Law, Probability & Risk 85 (2009); David H. Kaye, Probability, Individualization, and Uniqueness in Forensic Science Evidence: Listening to the Academies, 75 Brooklyn L. Rev. 1163 (2010).

Saturday, February 13, 2016

Broken Glass: What Do the Data Show?

In Broken Glass, Mangled Statistics, I noted "a plethora of statistical issues" in ASTM E2926-13, a Standard Test Method for Forensic Comparison of Glass Using Micro X-ray Fluorescence (μ-XRF) Spectrometry, that is working its way through the process for inclusion on the OSAC Registry of Approved Standards. The questions I raised about the Standard's procedures and criteria for declaring matches between glass specimens were based on elementary statistical theory but not data. Even if the ASTM's hypothesis testing procedures are idiosyncratic or conceptually flawed, they could have desirable properties.

There are some collections of glass that have been used to test the performance of the matching rules for some of the variables used in forensic testing. An FBI publication from 2009 offers the following summary:
Databases of refractive indices and/or chemical compositions of glass received in casework have been established by a number of crime laboratories (Koons et al. 1991). Although these glass databases are undeniably valuable, it should be noted that they may not be representative of the actual population of glass, and the distribution of glass properties may not be normal. Although these are not direct indicators of the rarity in any specific case, they can be used to show that the probability of a coincidental match is rare.

Koons and Buscaglia (1999) used the data from a chemical composition database and refractive index database to calculate the probability of a coincidental match. They estimated that ... the chance of finding a coincidental match in forensic glass casework using refractive index and chemical composition alone is 1 in 100,000 to 1 in 10 trillion, which strongly supports the supposition that glass fragments recovered from an item of evidence and a broken object with indistinguishable [refractive index] and chemical composition are unlikely to be from another source and can be used reliably to assist in reconstructing the events of a crime.

Range overlap on glass analytical data that include chemical composition data is considered a conservative standard. In one study, on a data set consisting of three replicate measurements each for 209 specimens, the range-overlap test discriminated all specimens, and all other statistical analysis-based tests performed worse (Koons and Buscaglia 2002).

Range-overlap tests, however, may achieve their high discrimination by indicating that two specimens from the same source are differentiable. Another study showed that when using a range-overlap test, the number of specimens differentiated that were actually from the same source may have been as high as seven percent (Bottrell et al. 2007).

The range-overlap approach, however, seems prudent given that other tests with higher thresholds for differentiation, such as t-tests with Welch modification (Curran et al. 2000) or Bayesian analysis (Walsh 1996), lower the number of specimens differentiated that were actually from the same source by worsening the ability to differentiate specimens that are genuinely different, a result that is unacceptable.
If I understand the argument, the author contends that high sensitivity is more important than high specificity. That makes sense for a screening test that will be followed by a more specific test, but in general, is it better to avoid falsely associating a defendant with crime-scene glass or to avoid falsely associating the defendant with the known glass? Any decision rule as to what is "indistinguishable" will generate a mix of false positives and false negatives. Should not the ASTM standards provide estimates from data (that might be representative of some relevant population) of these risks for each decision rule that the standards endorse or mandate?

References

Wednesday, February 3, 2016

Broken Glass, Mangled Statistics

The motto of ASTM International is “Helping Our World Work Better.” This internationally recognized standards development organization contributes to the world of forensic science by promulgating standards of various kinds for performing and interpreting chemical and other tests.

By mid-August 2015, five ASTM Standards were up for public comment to the Organization of Scientific Area Committees. OSAC “is part of an initiative by NIST and the Department of Justice to strengthen forensic science in the United States.” [1] Operating as “[a] collaborative body of more than 500 forensic science practitioners and other experts,” [1] OSAC is reviewing and developing documents for possible inclusion on a Registry of Approved Standards and a Registry of Approved Guidelines. 1/ NIST promises that “[a] standard or guideline that is posted on either Registry demonstrates that the methods it contains have been assessed to be valid by forensic practitioners, academic researchers, measurement scientists, and statisticians ... .” [2]

Last month, OSAC approved its first Registry entry (notwithstanding some puzzling language), ASTM E2329-14, a Standard Practice for Identification of Seized Drugs. Another standard on the list for OSAC’s quasi-governmental seal of approval is ASTM E2926-13, a Standard Test Method for Forensic Comparison of Glass Using Micro X-ray Fluorescence (μ-XRF) Spectrometry (available for a fee). It will be interesting to see whether this standard survives the scrutiny of measurement scientists and statisticians, for it raises a plethora of statistical issues.

What It is All About

Suppose that someone stole some bottles of beer and money from a bar, breaking a window to gain entry. A suspect’s clothing is found to contain four small glass fragments. Various tests are available to help determine whether the four fragments (the “questioned” specimens) came from the broken window (the “known”). The hypothesis that they did can be denoted H1, and the “null hypothesis” that they did not can be designated H0.

Micro X-ray Fluorescence (μ-XRF) Spectrometry involves bombarding a specimen with X-rays. The material then emits other X-rays at frequencies that are characteristic of the elements that compose it. In the words of the ASTM Standard, “[t]he characteristic X-rays emitted by the specimen are detected using an energy dispersive X-ray detector and displayed as a spectrum of energy versus intensity. Spectral and elemental ratio comparisons of the glass specimens are conducted for source discrimination or association.” Such “source discrimination” would be a conclusion that H0 is true; “association” would be a conclusion that H1 is true. The former finding would mean that the suspect's glass fragments did not come from the crime scene; the latter would mean either that (one way or another) they came from the broken window at the bar (or from another piece of glass somewhere that has a similar elemental composition).

Unspecified "Sampling Techniques"
for Assessing Variability Within the Pane of Window Glass

One statistical issue arises from the fact that the known glass is not perfectly homogeneous. Even if measurements of the ratios of the concentrations of different elements in a specimen are perfectly precise (the error of measurement is zero), a fragment from one location could have a different ratio than a fragment from another place in the known specimen. This natural variability must be accounted for in deciding between the two hypotheses. The Standard wisely cautions that “[a]ppropriate sampling techniques should be used to account for natural heterogeneity of the material, varying surface geometries, and potential critical depth effects.” But it gives no guidance at all as to what sampling techniques can accomplish this and how measurements that indicate spatial variation should be treated.

The Statistics of "Peak Identification"

The section of ASTM E-2926 on “Calculation and Interpretation of Results” advises analysts to “[c]ompare the spectra using peak identification, spectral comparisons, and peak intensity ratio comparisons.” First, “peak identification” means comparing “detected elements of the questioned and known glass spectra.” The Standard indicates that when “[r]eproducible differences” in the elements detected in the specimens are found, the analysis can cease and the null hypothesis H0 can be presented as the outcome of the test. No further analysis is required. The criterion for when an element “may be” detected is that “the area of a characteristic energy of an element has a signal-to-noise ratio of three or more.” Where did this statistical criterion come from? What is the sensitivity and specificity of a test for the presence of an element based on this criterion?

The Statistics of Spectral Comparisons

Second, “spectral comparisons should be conducted,” but apparently, only “[w]hen peak identification does not discriminate between the specimens.” This procedure amounts to eyeballing (or otherwise comparing?) “the spectral shapes and relative peak heights of the questioned and known glass specimen spectra.” But what is known about the performance of criminalists who undertake this pattern-matching task? Has their sensitivity and specificity been determined in controlled experiments, or are judgments accepted on the basis of self-described but incompletely validated “knowledge, skill, ability, experience, education, or training ... used in conjunction with professional judgment,” to use a stock phrase found in many an ASTM Standard?

The Statistics of Peak Intensity Ratios

Third, only “[w]hen evaluation of spectral shapes and relative peak heights do not discriminate between the specimens” does the Standard recommend that “peak intensity ratios should be calculated.” These “peak intensity ratio comparisons” for elements such as “Ca/Mg, Ca/Ti, Ca/Fe, Sr/Zr, Fe/Zr, and Ca/K” “may be used” “[w]hen the area of a characteristic energy peak of an element has a signal-to-noise ratio of ten or more.” To choose between “association” and “discrimination of the samples based on elemental ratios,” the Standard recommends, “when practical,” analyzing “a minimum of three replicates on each questioned specimen examined and nine replicates on known glass sources.” Inasmuch as the Standard emphasizes that “μ-XRF is a nondestructive elemental analysis technique” and “fragments usually do not require sample preparation,” it is not clear just when the analyst should be content with fewer than three replicate measurements—or why three and nine measurements provide a sufficient sampling to assess measurement variability in two sets of specimens, respectively.

Nevertheless, let’s assume that we have three measurements on each of the four questioned specimens and nine on the known specimen. What should be done with these two sets of numbers? The Standard first proposes a “range overlap” test. I’ll quote it in full:
For each elemental ratio, compare the range of the questioned specimen replicates to the range for the known specimen replicates. Because standard deviations are not calculated, this statistical measure does not directly address the confidence level of an association. If the ranges of one or more elements in the questioned and known specimens do not overlap, it may be concluded that the specimens are not from the same source.
Two problems are glaringly apparent. First, statisticians appreciate that the range is not a robust statistic. It is heavily influenced by any outliers. Second, if the  properties of the "ratio ranges" are unknown, how can one know what to conclude—and what to tell a judge, jury, or investigator about the strength of the conclusion? Would a careful criminalist who finds no range overlap have to quote or paraphrase the introduction to the Standard, and report that "the specimens are indistinguishable in all of these observed and measured properties," so that "the possibility that they originated from the same source of glass cannot be eliminated"? Would the criminalist have to add that there is no scientific basis for stating what the statistical significance of this inability to tell them apart is? Or could an expert rely on the Standard to say that by not eliminating the same-source possibility, the tests "conducted for source discrimination or association" came out in favor of association?

The Standard offers a cryptic alternative to the simplistic range method (without favoring one over the other and without mentioning any other statistical procedures):
±3s—For each elemental ratio, compare the average ratio for the questioned specimen to the average ratio for the known specimens ±3s. This range corresponds to 99.7 % of a normally distributed population. If, for one or more elements, the average ratio in the questioned specimen does not fall within the average ratio for the known specimens ±3s, it may be concluded that the samples are not from the same source.
The problems with this poorly written formulation of a frequentist hypothesis test are legion:

1. What "population" is "normally distributed"? Apparently, it is the measurements of the elemental ratios in the questioned specimen. What supports the assumption of normality?

2. What is "s"? The standard deviation of what variable? It appears to be the sample standard deviation of the nine measurements on the known specimen.

3. The Standard seems to contemplate a 99.7% confidence interval (CI) for the mean μ of the ratios in the known specimen. If the measurement error is normally distributed about μ, then the CI for μ is approximately the known specimen's sample mean ±4.3s. This margin of error is larger than ±3s because the population standard deviation σ is unknown and the sample mean therefore follows a t-distribution with eight degrees of freedom. The desired 99.7% is the coverage probability for a ±3σ CI. Using ±3 with the estimator s rather than the true value σ results in a confidence coefficient below 99%. One would have to use a number greater than ±4 rather than ±3 to achieve 99.7% confidence.

4. The use of any confidence interval for the sample mean of the measurements in the known specimen is misguided. Why ignore the variance in the measured ratios in the questioned specimens? That is, the recommendation tells the analyst to ask whether, for each ratio in each questioned specimen, the miscomputed 99.7% CI covers “the average ratio in the questioned specimen.” But this “average ratio” is not the true ratio. The usual procedure (assuming normality) would be a two-sample t-test of the difference between the mean ratio for the questioned sample and the mean for the known specimen.

5. Even with the correct test statistic and distribution, the many separate tests (one for each ratio Ca/Mg, Ca/Ti, Fe/Zr, etc.) cloud the interpretation of the significance of the difference in a pair of sample means. Moreover, with multiple unknown specimens, the probability of finding a significant difference in at least one ratio for at least one unknown fragment is greater than the significance probability in a single comparison. The risk of a false exclusion for, say, ten independent comparisons could be ten times the nominal value of 0.003.

6. Why ±3 as opposed to, say, ±4? I mention ±4 not because it is clearly better, but because it is the standard for making associations using a different test method (ASTM E2330). What explains the same standards development organization promulgating facially inconsistent statistical standards?

7. Why strive for a 0.003 false-rejection probability as opposed to, say, 0.01, 0.03, or anything else? This type of question can be asked about any sharp cutoff. Why is a difference of 2.99σ dismissed as not useful when 3σ is definitive? Within the classical hypothesis-testing framework, an acceptable answer would be that the line has to be drawn somewhere, and the 0.003 significance level is needed to protect against the risk of a false rejection of the null hypothesis in situations in which a false rejection would be very troublesome. Some statistics textbooks even motivate the choice of the less demanding but more conventional significance level of 0.05 by analogizing to a trial in which a false conviction is much more serious than a false acquittal.

Here, however, that logic cuts in the opposite direction. The null hypothesis H0 that should not be falsely rejected is that the two sets of measurements come from fragments that do not have a common source. But 0.003 is computed for the hypothesis H1 that the fragments all come from the same, known source. The significance test in ASTM E2926-13 addresses (in its own way) the difference in the means when sampling from the known specimen. Using a very demanding standard for rejecting H1 in favor of the suspect’s claim H0 privileges the prosecution claim that the fragments come from different sources. 2/ And it does so without mentioning the power of the test: What is probability of reporting that fragments are indistinguishable — that there is an association — when the fragments do come from different sources? Twenty years ago, when a National Academy of Sciences panel examined and approved the FBI's categorical rule of "match windows" for DNA testing, it discussed both operating characteristics of the procedure—the ability to declare a match for DNA samples from the same source (sensitivity) and the ability to declare a nonmatch for DNA samples from different sources. [3] By looking only to sensitivity, ASTM E2926-13 takes a huge step backwards.

8. Whatever significance level is desired, to be fair and balanced in its interpretation of the data, a laboratory that undertakes hypothesis tests should report the probability of differences in the test statistic as large or larger than those observed under the two hypotheses: (1) when the sets of measurements come from the same broken window (H1); and (2) when the sets of measurements come from different sources of glass in the area in which the suspect lives and travels (H0). The ASTM Standard completely ignores H0. Data on the distribution of the elemental composition of glass in the geographic area would be required to address it, and the Standard should at least gesture to how such data should be used. If such data are missing, the best the analyst can do is to report candidly that the questioned fragment might have come from the known glass or from any other glass with a similar set of elemental concentrations and, for completeness, to add that how often other glass like this is present is unknown.

9. Would a likelihood ratio be a better way to express the probative value of the data? Certainly, there is an argument to that effect in the legal and forensic science literature. [4-8] Quantifying and aggregating the spectral data that the ASTM Standard now divides into three, lexically ordered procedures and combining them with other tests on glass would be a challenge, but it merits thought. Should not the Standard explicitly acknowledge that reporting on the strength of the evidence rather than making categorical judgments is a respectable approach?

* * *

In sum, even within the framework of frequentist hypothesis testing, ASTM E2926 is plagued with problems — from the wrong test statistic and procedure for the specified level of “confidence,” to the reversal of the null and alternative hypotheses, to the failure to consider the power of the test. Can such a Standard be considered “valid by forensic practitioners, academic researchers, measurement scientists, and statisticians”?

Notes
  1. The difference between the two is not pellucid, since OSAC-approved standards can be a list of “shoulds” and guidelines can include “shalls.”
  2. The best defense I can think of for it is a quasi-Bayesian argument that by the time H1 gets to this hypothesis test, it has survived the qualitative "peak identification" and "spectral comparison" tests. Given this prior knowledge, it should require unusually surprising evidence from the peak intensity ratios to reject H1 in favor of the defense claim H0.
References
  1. OSAC Registry of Approved Standards and OSAC Registry of Approved Guidelines http://www.nist.gov/forensics/osac/osac-registries.cfm, last visited Feb. 2, 2016
  2. NIST, Organization of Scientific Area Committees, http://www.nist.gov/forensics/osac/index.cfm, last visited Feb. 2, 2016
  3. National Research Council Committee on Forensic DNA Science: An Update, The Evaluation of Forensic DNA Evidence (1996)
  4. Colin Aitken & Franco Taroni, Statistics and the Evaluation of Evidence for Forensic Science (2d ed. 2004)
  5. James M. Curran et al., Forensic Interpretation of Glass Evidence (2000)
  6. ENFSI Guideline for Evaluative Reporting in Forensic Science (2015)
  7. David H. Kaye et al., The New Wigmore: Expert Evidence (2d ed. 2011)
  8. Royal Statistical Soc'y Working Group on Statistics and the Law, Fundamentals of Probability and Statistical Evidence in Criminal Proceedings: Guidance for Judges, Lawyers, Forensic Scientists and Expert Witnesses (2010)
Postscript: See Broken Glass: What Do the Data Show?, Forensic Sci., Stat. & L., Feb. 13, 2016,

Disclosure and disclaimer: Although I am a member of the Legal Resource Committee of OSAC, the views expressed here are mine alone. They are not those of any organization. They are not necessarily shared by anyone inside (or outside) of NIST, OSAC, any SAC, any OSAC Task Force, or any OSAC Resource Committee.

Friday, January 29, 2016

The First OSAC-approved Standard for Forensic Science

"All results for every forensic science method should indicate the uncertainty in the measurements that are made, and studies must be conducted that enable the estimation of those values." Source: National Research Council Committee on Identifying the Needs of the Forensic Science Community, Strengthening Forensic Science in the United States: A Path Forward 184 (2009)

"It is expected that in the absence of unforeseen error, an appropriate analytical scheme effectively results in no uncertainty in reported identifications." Source: Standard Practice for Identification of Seized Drugs (ASTM E2329-14, § 4.2), added to the National Institute of Standards and Technology OSAC Registry of Approved Standards on Jan. 27, 2016

The response to comments from within OSAC on this and related text in the ASTM standard stated: "Editorial. The Seized Drug subcommittee intends to clarify the quoted language pertaining to uncertainty and error during the next ASTM revision of this document." E2329-14 Seized Drugs Response to LRC Comments FINAL.pdf (277K) SAC Chemistry/Instrument Analysis, Jan. 11, 2016.