Showing posts with label CODIS loci. Show all posts
Showing posts with label CODIS loci. Show all posts

Saturday, April 13, 2019

Questions on Kentucky's Rapid DNA Program for Sex Crime Investigations

An Associated Press report being picked up in newspapers such as the Washington Post and USA Today gives the impression that Kentucky is on the verge of replacing conventional DNA analysis of rape kits with a "rapid DNA" instrument. (1) The article states that
The equipment, known as the ANDE Rapid DNA system, can generate DNA identification from forensic samples in less than two hours, Kentucky law enforcement officials said. ...
“If you are a sexual predator in ... Kentucky, we’re going to come after you,” Kentucky State Police Commissioner Richard Sanders said at a press conference. “And we now have new equipment to come after you quicker. We now have another way to identify who you are.” ...
A 2017 federal law authorized the FBI director to issue standards and procedures for rapid DNA analysis, described as a fully automated process, according to an FBI website.The ANDE system received FBI approval last year for use in accredited laboratories, the company’s website says. A cheek swab is inserted into an ANDE device roughly the size of an office photocopier and results appear within hours. By comparison, DNA samples sent to conventional crime labs can take months to analyze.
The FBI has indeed approved the ANDE 6C Rapid DNA System for use in accredited laboratories -- but only with "[k]nown reference buccal DNA sample[s]." (2) In other words, an STR profile from a cheek swab from a known individual can be added to or checked against a database that is part of the Combined DNA Index System (CODIS). The profile from the male fraction of a sample in sexual assault case ascertained via a rapid DNA machine cannot be.
"The analysis of forensic samples by a Rapid DNA system is not ... permitted to be uploaded [or] searched in CODIS at this time." -- FBI (2)
So what does Kentucky propose to do with the rapid profiles? In a press release, Governor Matt Bevin enthused that they "can help us to identify an assailant in a matter of hours – allowing us to focus the investigations of sexual crimes more quickly than ever before." (3) Of course, it does not take months for a "conventional crime lab" to perform the steps that generate an electropherogram in the ordinary way and to interpret the data, but rapid DNA analysis is less labor intensive. That is a good thing, but how will Kentucky "identify an assailant in a matter of hours" from the rapid profile without using any CODIS database?

Even without a database, rapid DNA can help if there already is a suspect. It might further implicate the suspect, or it might exonerate him, propelling the investigation in a new direction. But that seems to be an afterthought in Kentucky's description of "how rapid DNA will be used." The press release (3) states that
By matching the DNA from the crime to an I.D. in a criminal database, this could give immediate data that focuses the investigation and get [sic] a rapist off the streets. By collecting an additional sample from the victim – just one more swab – a perpetrator’s identity can be matched within two hours."
How Rapid DNA will be used for sexual assault investigations:
● If a victim presents at a hospital or police station, and agrees to having DNA specimens taken for testing, a DNA ID can be generated to inform the investigation. This is a separate process from the Sexual Assault Kit but can be taken during the sexual assault exam.
● Once the male DNA is separated from the female DNA, an operator can test the male DNA for a definitive DNA Identification pattern. This process takes less than two hours.
● If a DNA ID is generated, it can be compared to criminal databases. If a match is found, this will provide important information to law enforcement about the attacker. It may exonerate an innocent suspect. It may connect other crimes to this attack.
The release does not explain how this "separate process" can "match[] the DNA from the crime to an I.D. in a criminal database" when the FBI has yet to approve rapid profiles for CODIS database searches. Another Kentucky State Police document (4) indicates that Kentucky somehow has been able to perform database searches with rapid DNA profiles:
The state of Kentucky has been running a pilot program for several months, using the ANDE Rapid DNA system to test samples from rape cases. It has proved invaluable in several cases where there is a match to DNA in a criminal database and to others in which a specific suspect was under consideration. This has convinced Kentucky officials to more broadly implement Rapid DNA testing.
To be clear, I have no objection to the use of a more efficient technology, but I am left puzzled. Has the FBI given Kentucky a special dispensation to use rapid profiles with CODIS databases? Is Kentucky using some database outside of that system?

And what is the basis for Kentucky's claim that inasmuch as "Rapid DNA uses Short Tandem Repeats for DNA ID, which has no coding information,.... there is no genealogy or health information gathered."? (3) The STRs certainly convey information on (close) genetic relationships. (5, 6) As for health-related information, they lie somewhere between none (like a passport number) and not much (like an ABO blood type). (7)

REFERENCES
  1. Bruce Schreiner  (AP), Kentucky To Use Rapid DNA Tests for Sex Assault Cases, Wash. Post, Apr. 10, 2019, https://www.washingtonpost.com/business/technology/kentucky-to-use-rapid-dna-tests-for-sex-assault-cases/2019/04/10/a28d484e-5bd4-11e9-98d4-844088d135f2_story.html
  2. Rapid DNA, www.fbi.gov/services/laboratory/biometric-analysis/codis/rapid-dna (viewed Apr. 13, 2019)
  3. Kentucky State Police, Kentucky Rape Reduction with ANDE Rapid DNA News Conference Fact Sheet, available via link at https://www.ande.com/kentucky-case-study/
  4. Kentucky State Police, The Use of Rapid DNA to End the Sexual Assault Epidemic, available via link at https://www.ande.com/kentucky-case-study/
  5. David H. Kaye, The Genealogy Detectives: A Constitutional Analysis of “Familial Searching”, 51 Am. Crim. L. Rev. 109 (2013), preprint available at ssrn.com/abstract=2043091
  6. Henry T. Greely & David H. Kaye, A Brief of Genetics, Genomics and Forensic Science Researchers in Maryland v. King, 53 Jurimetrics J. 43 (2013), available at ssrn.com/abstract=2403063
  7. David H. Kaye, Mopping Up After Coming Clean About "Junk DNA", Nov. 23, 2007, available at http://ssrn.com/abstract=1032094.

Thursday, July 5, 2018

A Strange Report of "Forensic Epigenetics ... in CODIS"

The “Featured Story” in today’s Forensic Magazine is “Forensic Epigenetics: How Do You Sort Out Age, Smoking in CODIS?” The obvious answer is that you don't and you can't. CODIS records contain no epigenetic data.

What Is Epigenetics?

As a Nature educational webpage explains, “[e]pigenetics involves genetic control by factors other than an individual's DNA sequence. Epigenetic changes can switch genes on or off and determine which proteins are transcribed.” 1/ One chemical mechanism for accomplishing this is DNA methylation, "a chemical process that adds a methyl group to DNA." 2/ More precisely, "methylation of DNA (not to be confused with histone methylation) is a common epigenetic signaling tool that cells use to lock genes in the 'off' position." 3/ This methylation is involved in cell differentiation and hence the formation and maintenance of different tissue types. 4/  "Given the many processes in which methylation plays a part, it is perhaps not surprising that researchers have also linked errors in methylation to a variety of devastating consequences, including several human diseases.” 5/

In forensic genetics, "DNA methylation profiling [has been proposed] for tissue determination, age prediction, and differentiation between monozygotic twins." 6/ Because this "profiling" can uncover health-related and other information as well, discussion of regulating its use by police has begun. 7/

What Is CODIS?

CODIS is “the acronym for the Combined DNA Index System and is the generic term used to describe the FBI’s program of support for criminal justice DNA databases as well as the software used to run these databases.” 8/ The DNA data, which come from twenty locations (loci) on various chromosomes, reveal nothing about methylation patterns. The information from these loci relates solely to the set of underlying DNA sequences. These particular sequences are not transcribed, are essentially identical in all tissues and all identical twins, and do not change as a person ages (except for occasional mutations).

What Is “Sort[ing] Out Age, Smoking in CODIS?”

I don't know. Forensic epigenetics or epigenomics involves neither CODIS databases, CODIS loci, nor CODIS software. Does Forensic Magazine's “Senior Science Writer” think that the databases will be expanded to include epigenetic data? That is not what the article asserts. The only attempt to bridge the two is a concluding sentence that reads, "But some studies, like a Stanford exploration last spring, show that even 13 loci can carry more information than originally believed."

That is not much of a connection, and the statement itself is a trifle misleading. Thirteen is the number of STR loci in CODIS profiles before the expansion to twenty in 2017. The description of the “Stanford exploration” referenced in the article 9/ does not show that the original understanding of the information contained in those core CODIS loci was faulty. Rather, it talks about the growth of the size of the databases and research showing that CODIS profiles “could possibly” be linked to records in medical research databases by “authorized or unauthorized analysts equipped with two datasets, one with SNP genotypes and another CODIS genotypes.” 10/

This possibility does not come as a complete surprise. CODIS profiles are meant to be individual identifiers (or nearly so). If there are genomic databases that sufficiently overlap these regions, then a CODIS profile can be used to locate the record pertaining to the same individual in those databases. The extent to which this possibility is cause for concern is worth considering, 11/ but it has nothing to do with the privacy implications of epigenetic data.

NOTES
  1. Simmons, D. (2008) Epigenetic influence and disease. Nature Education 1(1):6
  2. Id.
  3. Theresa Phillips (2008) The role of methylation in gene expression. Nature Education 1(1):116.
  4. See, e.g., Karyn L. Sheaffer, Rinho Kim, Reina Aoki, et al. (2014) DNA methylation is required for the control of stem cell differentiation in the small intestine. Genes & Development, http://genesdev.cshlp.org/content/28/6/652.abstract; Bo Zhang, Yan Zhou, Nan Lin, et al. (2013) Functional DNA methylation differences between tissues, cell types, and across individuals discovered using the M&M algorithm. Genome Research, https://genome.cshlp.org/content/early/2013/06/26/gr.156539.113.abstract
  5. Phillips, supra note 3.
  6. Athina Vidaki & Manfred Kayser (2017) From forensic epigenetics to forensic epigenomics: broadening DNA investigative intelligence, Genome Biol. 18: 238, doi:  10.1186/s13059-017-1373-1
  7. Mahsa Shabani, Pascal Borry, Inge Smeers, & Bram Bekaert (2018) Forensic Epigenetic Age Estimation and Beyond: Ethical and Legal Considerations. Trends in Genet 34(7): 489–491
  8. FBI, Frequently Asked Questions on CODIS and NDIS, https://www.fbi.gov/services/laboratory/biometric-analysis/codis/codis-and-ndis-fact-sheet
  9. Seth Augenstein, CODIS Has More ID Information than Believed, Scientists Find,” Forensic Mag., May 15, 2017, https://www.forensicmag.com/news/2017/05/codis-has-more-id-information-believed-scientists-find
  10. Id.
  11. Cf. David H. Kaye, The Genealogy Detectives: A Constitutional Analysis of “Familial Searching,” 51 Am. Crim. L. Rev. 109, 137 n. 170 (2013) (“However, there is at least one rather roundabout way in which the identification profiles could reveal substantial medical information. In the future, when the full genomes of individuals are recorded in clinical databases of medical records, a police agency possessing the profile and having surreptitious access to the database could locate the entry for the individual’s genome and any associated medical records without anyone’s knowledge. Although the STRs would be useful only for identification, that use could be the key to locating information in patient records. Furthermore, the patient’s records and full genome could lead police to the stored genomes and records of relatives. Although I cannot think of many scenarios in which police would be motivated to engage in this computer hacking and medical snooping, there may be some.”).

Tuesday, May 16, 2017

The Reappearing Rapid DNA Act

With bipartisan sponsorship, the Rapid DNA Act of 2017 (H.R.510 and S. 139) is sailing through Congress. The Senate bill made it to the legislative calendar on May 11, 2017, without amendment and without a written report from the Judiciary Committee.  The Committee Chairman, Senator Grassley, wrote this about the bill:
Turning to legislation, the first bill is S.139, the Rapid DNA Act of 2017. It is sponsored by Senator Hatch. The Committee reported this bill and the Senate passed it in the last Congress. The bill would establish standards for a new category of DNA samples that can be taken more quickly and then uploaded to our national DNA index. 1/
This characterization is misleading. The bill itself contains no standards for producing profiles to upload to the national database. It orders the FBI to “issue standards.” Specifically, the part of the bill entitled “standards” adds to the DNA Identification Act of 1994, 42 U.S.C. § 14131(a), a new Section 5, which reads as follows:
(A) ... the Director of the Federal Bureau of Investigation shall issue standards and procedures for the use of Rapid DNA instruments and resulting DNA analyses.
(B) In this Act, the term ‘Rapid DNA instruments’ means instrumentation that carries out a fully automated process to derive a DNA analysis from a DNA sample. 2/
But the FBI does not need new authorization to devise standards for “Rapid DNA instruments.” The “resulting DNA analyses” are not a new category of “samples,” and some such profiles already may be in the National DNA Index System (NDIS). In fact, the FBI issued standards for “rapid” profiles years ago. One need only peek at the FBI's forthright answers to “Frequently Asked Questions on Rapid DNA Analysis.” There, the FBI explained that
Based upon recommendations from the Scientific Working Group on DNA Analysis Methods (SWGDAM), the FBI Director approved and issued The Addendum to the Quality Assurance Standards for DNA Databasing Laboratories performing Rapid DNA Analysis and Modified Rapid DNA Analysis Using a Rapid DNA Instrument (or “Rapid QAS Addendum”). The Addendum contains the quality assurance standards specific to the use of a Rapid DNA instrument by an accredited laboratory; it took effect December 1, 2014.
The FBI added that “[a]n accredited laboratory participating in NDIS may use CODIS to upload authorized known reference DNA profiles developed with a Rapid DNA instrument performing Modified Rapid DNA Analysis to NDIS if [certain] requirements are satisfied” and that “DNA records generated by an NDIS-approved Rapid DNA system performing Rapid DNA analysis in an NDIS participating laboratory are eligible for NDIS.” 3/

But if the FBI does not need the bill to develop standards or to incorporate rapid-DNA results into NDIS, what is the real purpose of the bill? The answer is simple. The bill clears the way for these results to come, not from accredited laboratories, 4/ but from police stations, jails, or prisons. The House Judiciary Committee was explicit in its brief report on the bill:
Currently, booking stations have to send their DNA samples off to state labs and wait weeks for the results. This has created a backlog that impacts all criminal investigations using forensics, not just forensics used for identification purposes. H.R. 510 would modify the current law regarding DNA testing and access to CODIS. The short turnaround time resulting from increased use of Rapid DNA technology would help to quickly eliminate potential suspects, capture those who have committed a previous crime and left DNA evidence, as well as free up current DNA profilers to do advanced forensic DNA analysis, such as crime scene analysis and rape-kits. 5/
The FBI was more succinct when it referred to “the goal of using Rapid DNA systems in the booking environment” and reported that “legislation will be needed in order for DNA records that are generated by Rapid DNA systems outside an accredited laboratory to be uploaded to NDIS.6/

Is the migration of DNA profiling from the laboratory to the police station — and potentially to the officer on the street — a good idea? The efficiency argument from the House Committee has some force. We do not demand that only accredited laboratories conduct breath alcohol testing of drivers who seem to be intoxicated. Police using properly maintained portable instruments can do the job. 7/

How is DNA different? In one respect, it is less problematic than roadside alcohol testing. Rapid DNA analysis is not for crime-scene samples. (At least, not yet.) It is for samples from arrestees or convicted offenders whose profiles can be uploaded to a database. The police have an incentive to avoid uploading inaccurate profiles. Such profiles will degrade the effectiveness of the database. Any cold hits that they might produce will be shown to be false when a later DNA test from the suspect fails to replicate the incorrect profile. In contrast, incriminating output of a faulty alcohol test usually enables a conviction and will not be shown to be in error.

But there is more to the matter than efficiently generating and uploading profiles. It could be argued that DNA information is more private that a breath alcohol measurement and that having CODIS profiles known to local police is more dangerous than having it known only to laboratory personnel. Considering the limited kind of information that is present in a CODIS profile, however, this argument does not strike me as compelling.

POSTSCRIPT

The Rapid DNA Act of 2017 met no opposition as the Senate and House passed the bills. S. 139 generated unanimous consent (and no discussion) on May 16. 8/ Its counterpart, H.R. 510, passed after receiving praise from two of its sponsors and the observation from Representative Goodlatte (R-VA) that "this is a good bill. It is a bipartisan bill. I thank Members on both sides of the aisle for their contributions to this effort." 9/

NOTES
  1. Prepared Statement by Senator Chuck Grassley of Iowa, Chairman, Senate Judiciary Committee Executive Business Meeting, May 11, 2017, https://www.judiciary.senate.gov/imo/media/doc/05-11-17%20Grassley%20Statement.pdf, viewed May 16, 2017.
  2. Rapid DNA Act of 2017, S. 139 § 2(a).
  3. The difference between “Rapid DNA Analysis” and “Modified Rapid DNA Analysis” is that the former is “a “swab in – profile out” process ... of automated extraction, amplification, separation, detection, and allele calling without human intervention,” whereas the latter uses “human interpretation and technical review” for ascertaining the alleles in a profile. FBI, Frequently Asked Questions on Rapid DNA Analysis, https://www.fbi.gov/services/laboratory/biometric-analysis/codis/rapid-dna-analysis, Nos. 1 &2, viewed May 17, 2017.
  4. The DNA Identification Act of 1994, 42 U.S.C. § 14131, which the Rapid DNA Act amends, requires the FBI to create and consider the recommendations of "an advisory board on DNA quality assurance methods." § 14131(a)(1)(A).  The members of the board must come from "nominations proposed by the head of the National Academy of Sciences and professional societies of crime laboratory officials." Id. They "shall develop, and if appropriate, periodically revise, recommended standards for quality assurance, including standards for testing the proficiency of forensic laboratories, and forensic analysts, in conducting analyses of DNA." § 14131(a)(1)(C). As the name indicates, the board is purely advisory. The Act only demands that
    The Director of the Federal Bureau of Investigation, after taking into consideration such recommended standards, shall issue (and revise from time to time) standards for quality assurance, including standards for testing the proficiency of forensic laboratories, and forensic analysts, in conducting analyses of DNA.
    § 14131(a)(2).
    The advisory board was a half-a-loaf response to the recommendation of a National  Academy of Sciences committee for "a National Committee on Forensic DNA Typing (NCFDT) under the auspices of an appropriate government agency, such as NIH or NIST, to provide expert advice primarily on scientific and technical issues concerning forensic DNA typing." NRC Committee on DNA Technology in Forensic Science, DNA Technology in Forensic Science 72-73 (1992). Now that NIST has established an Organization of Scientific Area Committees for Forensic Science to develop science-based standards for DNA testing and other forensic science methods, Congress should reconsider the need for the overlapping FBI board.
  5. On May 11, 2017, the House Committee on the Judiciary recommended adoption of H.R. 510 without holding hearings. The Judiciary Committee saw no need to consult independent scientists. It was satisfied with the fact that
    the Judiciary Committee’s Subcommittee on Crime, Terrorism, Homeland Security and Investigations held a hearing on a virtually identical bill, H.R. 320, on June 18, 2015, [at which] testimony was received from: Ms. Amy Hess, Executive Assistant Director of Science and Technology, Federal Bureau of Investigation; Ms. Jody Wolf, Assistant Crime Laboratory Administrator, Phoenix Police Department Crime Laboratory, President, American Society of Criminal Laboratory Directors; and Ms. Natasha Alexenko, Founder, Natasha’s Justice Project.
    Report to accompany H.R. 510, May 11, 2017, https://www.congress.gov/115/crpt/hrpt117/CRPT-115hrpt117.pdf
  6. FBI Answers, No. 13, https://www.fbi.gov/services/laboratory/biometric-analysis/codis/rapid-dna-analysis, viewed May 17, 2017 (emphasis added).
  7. “As of January 1, 2017, there is no Rapid DNA system that is approved for use by an accredited forensic laboratory for performing Rapid DNA Analysis.” Several systems had been approved but they do “not contain the 20 CODIS Core Loci required as of January 1, 2017.” FBI Answers, No. 6, https://www.fbi.gov/services/laboratory/biometric-analysis/codis/rapid-dna-analysis, viewed May 16, 2017. 
  8. 163 Cong. Rec. S2954-2955, 115th Cong., 1st Sess., May 16, 2017.
  9. Id. at H4205.

Sunday, January 4, 2015

Buza Reloaded: California Balancing

This is the fourth installment on Buza II, the opinion of the California court of appeal that invalidates the state's DNA-on-arrest law. It discusses the part of the opinion that argues that the balance the U.S. Supreme Court struck in Maryland v. King is either flatly wrong or wrong for California. In giving substantial weight to concerns over "familial searching" and the information content of DNA samples, the opinion assumes that it is appropriate to strike down a law that is constitutionally reasonable as currently implemented because future developments might make it unreasonable as then implemented. This premise is highly contestable.

Formally, the conclusion that California's DNA-BC (Before Conviction) law is unreasonable under the Fourth Amendment as it appears in the California Constitution does not imply that it is unreasonable under the Fourth Amendment as it exists in the U.S. Constitution. California is a sovereign state of the Union, and its courts can read different meanings into the words of its constitution. But many of the reasons the Buza II opinion gives for its conclusion—if correct—also apply to nearly all of the 25 or so DNA-BC laws on the books, and the opinion itself indicates that, in large part, the divergence between Buza II and King emanates from the California judges’ outright disagreement with the Supreme Court's balancing in King.

To begin with, the California judges complain that King “unjustifiably dismissed concerns about the extent of the personal information contained in DNA samples by limiting ... attention to the profile used in DNA databanks, as currently restricted by statutes and scientific capability.” One might expect that this observation immediately would be followed by the undeniable fact that the entirety of a person’s genome contains some medically significant information that would not otherwise be known, such as predispositions to certain diseases. Testing for these alleles (or for markers for them) would pose significant privacy issues (which is why such testing generally is prohibited without the individual’s consent).

But the opinion veers off into a superficial discussion about the CODIS profile itself. The problem, according to Buza II, is that the profile can be used not merely to identify an individual whose DNA is taken when he is arrested, but also sometimes can be used to identify a first-degree relative as a likely source (when the arrestee’s DNA is a close mismatch to the crime-scene sample). This “familial searching,” as the court calls it, is a “factor not relevant to identity,” and therefore “present[s] additional privacy concerns.”

The second part of this statement is true enough. Like a perfect match, a close mismatch is relevant to the identity of the DNA source, but it also reveals that the arrestee could be genetically related to the source of the crime-scene DNA. 1/ Consider the “Grim Sleeper” case of serial rapes and murders in the Los Angeles area, with years of apparent inactivity between some of the attacks. Trawls of the database proved fruitless—until Christopher Franklin was convicted of a felony. His DNA profile did not match the Grim Sleeper’s, but it lined up with it in a manner that would be expected if the two were father and son. This led investigators to Christopher’s father, Lonnie Franklin, Jr. In this way, Lonnie emerged as a suspect only because of his son’s conviction. (His DNA profile was not in the database because his arrests had occurred before California had a database.) Now he stands accused of ten murders.

People v. Franklin reveals an important fact about kinship trawling. In Franklin, it is difficult to discern the slightest “additional privacy concerns.” That Lonnie was Christopher’s father was a publicly known fact, not a private secret. Furthermore, Lonnie can hardly claim to have a legitimate Fourth Amendment interest in keeping secret the fact that it was his DNA that was found on or around murdered women. 

Of course, there could be other cases in which the familial relationship between the database inhabitant and the culprit was not known to one or both of the genetically related individuals. In such situations, the claim to a right to keep the genetic relationship secret is more plausible. But the existence of possible cases of this kind does not demonstrate that the occasional legitimate privacy interests that might be affected by the rare, "other-directed" trawls (that look for people outside of the database) outweigh those of the government.

In particular, for Mark Buza and his relatives to have an additional privacy interest compromised by the arresteee database, at least two conditions would have to be fulfilled. First, California would have to initiate other-directed trawls of its arrestee database. It has never done so, and it cannot do so under the policy its Department of Justice has adopted for such database trawling. This policy confines the other-directed trawling to convicted-offender databases. Second, Mark Buza would have to have publicly unknown first-degree relatives whose DNA profile would be close enough to Mark’s to implicate them in other crimes via a kinship match to Mark’s profile.

On its face, the first condition suggests that the parts of the opinion discussing “familial searching” are inapposite. Why strike down a law because of what could be but is not? Nonetheless, the Buza court’s sensitivity to the possibility of a change in the state’s DNA-BC practice might be seen as prescient rather than premature. From the outset, an argument against DNA databases has been mission creep. Once the database is established, the state will be tempted to use it for additional and more insidious purposes. To guard against this outcome, the argument goes, society should bind itself to the mast in anticipation of an irresistible siren song.

There are situations in which this self-disabling strategy is advisable. Indeed, much of the Bill of Rights constrains the majority from doing what seems expedient or appealing in the heat of the political moment. But it is not so clear that a handful of judges should block the democratic decision to allow DNA-BC to be used in acceptable ways that advance law enforcement on the ground that the system might be administered in unacceptable ways at some future time. If and when a jurisdiction combines other-directed trawling and DNA-BC, courts can consider whether that type of trawling is so serious an invasion of privacy as to render it unconstitutional. Cf. United States v. Knotts, 460 U.S. 276 (1983) ("if such dragnet type law enforcement practices as respondent envisions should eventually occur, there will be time enough then to determine whether different constitutional principles may be applicable."). Using the mere possibility of a correctable change in the allowed uses of the DNA data to strike down the collection and otherwise acceptable uses of the data seems Draconian.

Moreover, relying on future familial searching as a ground for striking down the system as currently implemented is inconsistent with Buza II’s effort to distinguish the Maryland practice. Presiding Justice Kline emphasized the existence of a Maryland statute banning familial searching. But as Chief Judge Alex Kozinski of the U.S. Court of Appeals for the Ninth Judicial Circuit tartly observed in oral argument in Haskell v. Harris (a separate case challenging California DNA-BC law), statutes can be changed too. The logic of Buza II—that databases that are constitutionally reasonable (as currently implemented) but might become unreasonable (as implemented in the future) are constitutionally unreasonable ab initio—would render the Maryland law on DNA-BC unconstitutional.

Despite these problems, Buza II applies the nip-it-in-the-bud reasoning not only to DNA profiles but also to samples. Displaying little knowledge of behavioral genetics, the court invokes “the pedophile gene” and “the violence gene” that, it imagines, might well be discovered some day. It predicts that “surely law enforcement will seek to mine genetic information for that ‘identification purpose.’” 
But there is no good reason to believe that the word “identification” as used in DNA-BC laws would permit predictive genetic testing for these behaviors, and the court makes no attempt to explain why such testing could not be condemned as constitutionally unreasonable if and when the time arises.

My criticism of the court of appeal's reliance on dystopic visions of the future is not based on naive faith in the goodness of police and law enforcement laboratories. Courts need not—and should not—trust law enforcement to exercise perfect self-restraint in investigative methods that easily can be abused. Before approving a DNA database system, they should satisfy themselves that sufficient safeguards against predictable abuses are in place. But if such protections are present, courts should not invalidate a system because the safeguards might be removed or might cease to be effective in the future. In this case, might does not make the decision right.

Note
  1. Confusingly, the court presents this fact as if it "disproves the King majority’s assumption that 'the CODIS loci come from noncoding parts of the DNA that do not reveal the genetic traits of the arrestee.'" Some noncoding DNA does affect visible traits of an arrestee, but the CODIS loci, as far as current science can tell, do not reveal much about any phenotypes. Because all DNA sequences are inherited, however, including those that King (also confusingly) calls "junk," the ones that vary across individuals, can be used in kinship analysis. In fact, the sequences that do give rise to individual traits often are the best for this purpose because they tend to be extremely variable within populations.
References
Closely related postings

Thursday, January 3, 2013

"Scientists' Brief" on CODIS Loci: Q & A

On November 9, 2012, the Supreme Court voted to review a case posing the following question: “Does the Fourth Amendment allow the States to collect and analyze DNA from people arrested and charged with serious crimes?” In Maryland v. King, the state’s supreme court concluded that the protection against unreasonable searches and seizures forbids the state from collecting DNA from an individual whose true identity can be established with ordinary fingerprints. On December 28, 2012, the Supreme Court received a Brief of Genetics, Genomics and Forensic Science Researchers as Amici Curiae. Below are several questions and answers about the brief.

Who contributed to the brief?

I did, and Hank Greely was an additional author. The scientists who participated in the writing are all active and distinguished researchers at medical schools (including Harvard, Yale, and Johns Hopkins) or universities (including Duke, Penn State, and Kings College, London). They include a former president of the American Society of Human Genetics, a past president of the American Board of Medical Genetics, Fellows of the American Association for the Advancement of Science, and members of the Institute of Medicine and the American Academy of Arts and Sciences.

Why did these law professors, medical and statistical geneticists, and molecular biologists submit an amicus brief?

The brief is intended “to inform the Court of the possible medical and social significance of the DNA data stored in law enforcement databases.” (P. 1). Advocacy groups, legal scholars, and some judges have asserted that the small number of features used in law enforcement DNA databases are predictive of health status (or soon will be). The brief attempts to clarify this issue.

Which side does the brief support?

The brief was submitted in support of neither side. It describes the nature of genetic information, the features of the genome used in law enforcement DNA databases, how those features are used in medical research, and whether they currently permit police, employers, or insurers to discern significant facts about a person’s present or future health status.

What conclusions does it reach?

Amici conclude that “[u]nlike medical genetic tests, law enforcement identification profiles have no known value for medical diagnosis or prediction of future health.” (P. 2).

That’s today. What about the future?

Amici caution that “no one can say with certainty what the future will bring, and it is possible that specific loci will be found to affect the operation of certain genes or to display correlations to disease states.” (P. 2). Nevertheless, they suggest that “it is unlikely that the identification profiles will turn into powerful medical diagnostic or predictive tools that can be used to infer disease states or predispositions by examining forensic database records.” (P.2).

Does this mean that the “CODIS loci,” as the identifying features are called, have no medical significance?

Absolutely not. The DNA sequences have been used in medical research for some 20 years to hunt for disease-causing gene mutations. They have been studied for associations with diseases and traits such as longevity. The question the brief addresses is what kind of information can be gleaned from inspecting a database record.

Doesn’t the highly publicized ENCODE Project prove that there is no such thing as “junk DNA”?

The brief contends that debate over the fraction of the genome that is, in an evolutionary sense, 'junk' ... is orthogonal to the matter before the Court. (P. 26). A section of the brief explains that the data sets and papers recently released from the international Encyclopedia of DNA Elements Project are important to further research into gene regulation and other matters, but they do not indicate that all DNA sequences are critical to health or other important traits. What “[t]he ENCODE papers show [is] that 80% of the genome displays signs of certain types of biochemical activity—even though the activity may be insignificant, pointless, or unnecessary.” (P. 32).

Well, how about other uses? Don’t the CODIS loci tell scientists a lot about a person’s ancestry and race?

Not really. The CODIS loci can reveal something about bio-geographic ancestry, but anthropologists and population geneticists use far more probative ancestry-informative and lineage markers to study genetic histories. That “race” is not a biological category is now well known. As for socially perceived race, “[a] CODIS profile could be used to calculate probabilities that someone would be described as Caucasian, African-American, or Hispanic, but categorical inferences would not be very accurate, and attempts to predict the census-type race of a person from a CODIS profile would seem pointless considering that apparent race already would be known.” (P. 36).

So the brief shows that there is absolutely no important information that can be deduced from a CODIS profile?

No, amici do not say that either. The brief explains that “[b]ecause children inherit all their DNA from their biological parents, the CODIS loci can be powerful tools for determining whether two people could be genetically related as parent and child. ... [T]he most powerful genetic information other than identity that the CODIS profiles contain [would be] that two people are not parent and child” or “that two people were identical twins.” (Pp. 33-34).

Where can I find the brief?

Here is a pdf. It also should appear."soon," along with other briefs, on the American Bar Association's Preview of Supreme Court cases.

Postscript: The brief and an introduction to it is published as Henry T. Greely & David H. Kaye, A Brief of Genetics, Genomics and Forensic Science Researchers in Maryland v. King, 53 Jurimetrics J. 43 (2013). The publication is available at http://ssrn.com/abstract=2403063

Wednesday, July 25, 2012

CODIS Loci Not Ready for Disease Prediction After All?

Last month, I noted the findings of a superior court in Vermont that "some of the CODIS loci have associations with identifiable serious medical conditions," making the scientific evidence "sufficient to overcome the previously held belief[s]" about the innocuous nature of the CODIS loci [1]. The judge based her conclusion in State v. Abernathy [2] that the CODIS loci now permit "probabilistic predictions of disease" on the unpublished views of biologist Greg Wray, who oversees the Center for Evolutionary Genomics and the DNA Sequencing Core Facility, within Duke University’s Institute for Genome Sciences and Policy.

A technical report accepted for publication in the Journal of Forensic Sciences seems to dispute these claims. Sara Katsanis, a staff researcher at the same Institute for Genome Sciences and Policy, and Jennifer Wagner, a research associate at the University of Pennsylvania’s Center for the Integration of Genetic Healthcare Technologies, searched the biomedical literature and genomic databases not only for associations with phenotypes in the current 13 loci used in offender databases, but also in ones that soon may be added to the system. They came up with “no evidence” that any particular CODIS single-locus genotypes “are indicative of phenotype.”

References

1. CODIS Loci Ready for Disease Prediction, Vermont Court Says, June 15, 2012.
2. State v. Abernathy, No. 3599-9-11 (Vt. Super. Ct. June 1, 2012).

Friday, June 15, 2012

CODIS Loci Ready for Disease Prediction, Vermont Court Says

A trial court in Vermont has gone where no court has gone before. In State v. Abernathy [1], Chittenden Superior Court Judge Alison Sheppard Arms found that because "[s]ix CODIS loci ... have associations with an increased risk of disease or have functional properties," the custodians of law enforcement DNA databases can make "probabilistic predictions of disease." According to the judge, modern research has established that "some of the CODIS loci have associations with identifiable serious medical conditions," making the scientific evidence "sufficient to overcome the previously held belief[s]" about the innocuous nature of the CODIS loci.

Emphasizing this finding that richly information-laden STR profiles reside in identification databases, the court proceeded to strike down "Vermont's new pre-conviction DNA testing requirement ... that requires submission of a DNA sample from a 'person for whom the court has determined at arraignment there is probable cause that the person has committed a felony ... .'" In an atypical opinion, the court applied a "special needs" balancing test, placed the burden of proof on the state, and held that this law violates the state constitution.

A major theme in Judge Arms' discussion of human genetics is that there has been a revolution in our understanding of what used to be called "junk DNA." Even though the CODIS loci originally were described as "junk" in "good faith," that understanding was wrong--we now know that even DNA that does not code for proteins is biologically important.1 Other judges, advocacy groups, and at least one law professor have jumped from the discovery that the triplet code for proteins is not the sole message inscribed in DNA to the conclusion that all the CODIS loci may well convey significant information about disease states or propensities.

There are a couple of problems with this reasoning. All that we actually know is that some non-protein-coding DNA regulates gene expression. Scientists do not believe that all non-protein-coding sequences are regulatory. In particular, whether noncoding, nontranscribed, and largely nonconserved sequences are part of a regulatory system (even if their presence might have some function) is far from established.2 The opinion cites an essay I wrote making this point [3] but then ignores its content. It quotes the legal treatise, Modern Scientific Evidence, for the view that "while it is generally agreed that no single loci [sic] contains a gene that definitively determines any discernible characteristic of significance, there are nonetheless indications that they may play a role in some sensitive matters, and continued debates about their importance." Before Abernathy, it appeared that the "continued debates" ended five years ago with agreement with what already was clear -- that even if the loci do not play a functional role, they might, like certain fingerprint patterns or blood types, have some statistical associations with diseases.3

Venturing beyond the inconclusive generalities like these, Abernathy refers to the biomedical literature on five loci and to a testifying expert's characterization of the literature (with no specific references) on another locus. The opinion does not give the magnitude of any putative association, let alone any measure of predictive utility.4 It uses the following phrases: "a fairly large effect size," "a modest association," "not the most strongly associated," "small but ... not zero,"5 and "cannot find that this marker has no association." It does not provide measures of the uncertainty in these estimates. Finally, the opinion does not discuss the extent to which the studies said to show that the associations are real have been replicated.6

Of course, few judges could confidently review the flood of studies on human genetics. Unlike some previous opinions and law review articles, however, this opinion does not rely entirely or largely on newspaper headlines and stories about "junk DNA." Here, the iconoclastic findings came after an evidentiary hearing. But, as has happened before with DNA evidence [8], the evidentiary hearing was one-sided. The defendants presented the testimony of Professor Gregory Wray of Duke University, a specialist in genetics and evolutionary biology, and the state chose not to present an expert in medical genetics or genomics to counter his testimony. Although Professor Wray reviewed the biomedical literature before he testified, the defense submitted no written report, and the state rather than the defense introduced the papers cited in the opinion as exhibits. Scanning the testimony, it seems to me that Dr. Wray never was asked a series of critical questions:
  1. Is it generally accepted that the associations he pointed to apply to the population of individuals whose DNA is placed in law enforcement databanks?
  2. Assuming that they do apply to that population, what is the positive and negative predictive value of any inferences about disease based on the CODIS alleles?
  3. How would the predictive or diagnostic disease-related information in a state DNA database compare to that of (a) color photographs, (b) fingerprints, (c) the blood types used in old-fashioned serology, and (d) the HLA-A and HLA-B haplotypes once used on parentage testing?
  4. Are the CODIS genotypes likely to be substantially more predictive in the future?
Until these questions are answered, there is reason to ask whether the trial court's findings fairly represent the scientific status quo or instead are grim predictions of what could come to pass.

Notes

1. For a short audio clip reporting on the revolutionary discoveries, click on Joe Palca, Don't Throw It Out: 'Junk DNA' Essential In Evolution, All Things Considered, Aug. 19, 2011 (with a sound bite from Professor Gregory Wray, among other interviewees).

2. According to Judge Arms,"[t]he term 'junk DNA' was coined in the early 1980s." In fact, the phrase normally is attributed to Susumu Ohno, who used it in the title of a 1972 paper [2]. Ohno did not reason that "we don't know what noncoding DNA does, therefore, is it is useless junk." Indeed, he proposed that the duplication and inactivation of genes produce non-protein-coding DNA (now designated pseudogenes) that might have a function. A video introducing Ohno and reading an excerpt from the paper about the role of the noncoding sequences as "spacers" with evolutionary importance can be found at http://www.youtube.com/watch?v=nomI35DJB40&noredirect=1. Since 1972, other possible functions for noncoding DNA have been proposed. Some functions imply that the sequences should be conserved as one species evolves into another. Others, such as Ohno's suggestion that noncoding sequences act as buffers between genes, do not.

3. See [5, p. 228] (referring to "a brief debate in the legal literature" necessitated by "a misunderstanding by Simon Cole over some of the things I [John Butler] had written in a review article on STR markers" and emphasizing that "STR markers used for human identity testing do not predict disease."). One source of confusion, which also infects the Abernathy opinion is the thought that a statistical association between a locus and a disease detected in a family study in say, Northern India, establishes that the same association exists throughout the United States population.

4. Even a strong association (large relative risk) would not make for a useful predictive test if the prevalence of the condition is very small. See [3].

5. The sentence "[t]he relative risk of developing schizophrenia associated with this marker is small but it is not zero" is technically flawed. A relative risk of 1 would express a 0 correlation.

6. Replication is always important, and the problem of false positives is especially acute with genome-wide association studies. See, e.g., [6, 7].

References

1. State v. Abernathy, No. 3599-9-11 (Vt. Super. Ct. June 1, 2012).

2. S. Ohno, So Much "Junk" DNA in our Genome, 23 Brookhaven Symp. Biol. 366 (1972) (also published in Evolution of Genetic Systems 366 (H.H. Smith ed. 1972).

3. David H. Kaye, Please, Let's Bury the Junk: The CODIS Loci and the Revelation of Private Information, 102 Nw. U. L. Rev. Colloquy 70 (2007).

4. David H. Kaye, Mopping Up After Coming Clean About "Junk DNA", Nov. 23, 2007, available at http://ssrn.com/abstract=1032094.

5. John M. Butler, Advanced Topics in Forensic DNA Typing: Methodology (2012).

6. D.J. Hunter & P. Kraft, Drinking from the Fire Hose--Statistical Issues in Genomewide Association Studies, 357 N. Engl. J. Med. 436 (2007).

7. Thomas A. Pearson, & Teri A. Manolio, How to Interpret a Genome-wide Association Study, 299 J. Am. Med. Ass'n 1335 (2008).

8. David H. Kaye, The Double Helix and the Law of Evidence (2010).

Cross-posted from The Double Helix Law Blog.

Saturday, July 9, 2011

Junk Science in United States v. Pool

Having granted en banc review in United States v. Pool, 621 F.3d 1213 (9th Cir. 2010), the U.S. Court of Appeals for the Ninth Judicial Circuit is likely to produce as wild a set of conflicting opinions on DNA databases as it did in United States v. Kincade, 379 F.3d 813 (9th Cir. 2004).

The panel that heard the Pool case divided 2-1 and generated 3 opinions. Judge Callahan wrote an opinion upholding the federal law on taking DNA after an arrest. Visiting Judge Lucero (of the Tenth Circuit) joined in this opinion, but he also wrote a separate opinion. Judge Schroeder dissented.

The briefs that informed these three opinions left something to be desired. Here, I'll focus on one of my pet peeves--disingenuous or inane claims about the CODIS STR loci as a threat to privacy.

Appellant's Opening brief (available from a link on EPIC's website, along with a one-sided list of vaguely related articles) is rather shameless in this regard. It starts with the claim that "DNA profiles derived by STR may yield probabilistic evidence of the contributor’s race or sex." [1] Probabilistic evidence of sex from autosomal STRs? The arresting officers or jailers need a genetic test for that?

Then the brief cites Simon Cole's writing to support its sweeping statement that "scientific studies have debunked the notion that these regions of the genetic code are devoid of any biological function." Yet, the brief cites no study that "debunks" the notion that the length polymorphisms of the CODIS tetranucleotide STRs lack "biological function." The concurring opinion of Judge Lucero recognizes that Cole rejects the claim of functionality (for the moment). [2] However, a group in France has a theory and some data for a mechanism through which one such STR could regulate the expression of an enzyme. [3]

Finally, the brief proposes that the "specter of discrimination and stigma could arise where one or more STRs is found to correlate with another genetic marker whose function is known, so that the presence of the seemingly innocuous STR serves as a 'flag' for that genetic predisposition or trait." [4] An accompanying footnote gives this example: "A study in England from 2000 found that one of the markers used in DNA identification is closely related to the gene that codes for insulin, which itself relates to diabetes." [5]

The accused STR is TH01. It has been used in many studies investigating the association between (a) SNPs, VNTRs, and this STR in a complex of genes and (b) a large number of diseases. Unsurprisingly, associations have been observed. Some of the reported associations were spurious and were not replicated. Other associations probably are real. This does not mean that TH01, by itself, is a useful predictor of any of these diseases in a given population. In fact, one forensic biologist used the 2000 paper cited in Pool's brief to show that "such associations [between forensic STRs and disease-causing alleles in genes] are so ridiculously weak that serious protest could never form." [6] His explanation follows:
This is illustrated well by the possible association between certain alleles of an STR named TH01 and diabetes type 1 (Bennett and Todd, 1996; Stead et al., 2000). TH01 alleles are used routinely in DNA typing, and for a minute, the manufacturers of genetic fingerprint kits started to feel the heat over the possible association between an exonic illness and an intronic allele. Fortunately, it takes just a pen and a piece of paper to brush off possible concerns: four out of 1000 Europeans will eventually get diabetes type 1. If you carry one of the ‘risk’ alleles in the intronic TH01 region, your chances of getting diabetes type 1 is 0.13 out of 1000. If I find out that you are carrying the alleged risk allele in my laboratory during DNA typing, I could—but I am not allowed to—calculate your total risk for diabetes as 0.4 × 1.3 = 0.52%. In plain language: in the worst case scenario, one allele of your possible genetic fingerprint might tell me that your general risk of getting diabetes type 1 is increased from 0.4 to 0.52%. All other alleles will not tell me anything about you, or your potential risk for illnesses. Abuse of such information is impossible because it simply has no practical predictive value.
I do not want to "brush off possible concerns," and I understand the pressures and temptations of advocacy. Still, I wonder whether the Sacramento Federal Defender consulted the scientific literature on TH01 before citing an old article. Or whether he knew that the claims in the law review essay cited in the brief were the subject of an extensive rejoinder in the same journal. [7] If he did, he choose not to share this fact with the court. To my mind, that is not good advocacy.

Notes

1. Brief at 12 (quoting from a plurality opinion in Kincade).
2. 621 F.3d at 1230.
3. See Rolando Meloni, Post-genomic Era and Gene Discovery for Psychiatric Diseases: There Is a New Art of the Trade? The Example of the HUMTH01 Microsatellite in the Tyrosine Hydroxylase Gene, 26 Molecular Neurobiology 389 (2001).
4. Brief at 12.
5. Id. at 12 n.8.
6. Mark Benecke, Coding orNon-coding?, That Is the Question, 3 European Molecular Biology Organization Reports 498 (2002).
7. David H. Kaye, Please, Let's Bury the Junk: The CODIS Loci and the Revelation of Private Information, 102 Nw. U. L. Rev. Colloquy 70 (2007).

Crossposted from the Double Helix Law Blog