Saturday, March 21, 2015

The Junk DNA Wars

This month, the New York Times' published a report on “the junk DNA wars” asking “Is Most of Our DNA Garbage”? 1/ Readers of the article (and an anonymous follow-up piece on the reactions appearing in science blogs) 2/ would come away thinking that there is a serious debate in the scientific community over the proposition that “junk DNA” is “mostly functional.”

Without defining terms like “functional” and “junk,” however, it is impossible to know what is in dispute and what is not.The follow-up piece is particularly frustrating. It observes that
Some scientists, like T. Ryan Gregory, a evolutionary biologist ... argue that if DNA is mostly functional, then it’s hard to explain why rather humble species, like the onion, have far more DNA than we do. ...
Those who disputed Gregory’s findings [sic — Gregory did not discover the long-standing C-value paradox 3/ ], including supporters of intelligent design, cited the Encode Project, an N.I.H.-sponsored attempt to catalog the functional elements of the genome. Encode scientists found that 80 percent of the genome had “biochemical functions,” suggesting that there was a lot less junk DNA than scientists had thought. But did “biochemical function” really mean anything?
For many scientists, it didn’t. A University of Toronto biochemist, Larry Moran, wrote that “the general public has been snowed by the Encode publicity campaign and by naïve journalists who have enthusiastically reported that junk DNA is dead.”
But the Times' writers did not explain why “many scientists” are not snowed by the 80% statistic. After reading some of the ENCODE papers and the surrounding (typically hyperbolic) publicity, I concluded that:
The ENCODE papers show that 80% of the genome displays signs of certain types of biochemical activity—even though the activity may be insignificant, pointless, or unnecessary. This 80% includes all of the introns, for they are active in the production of pre-mRNA transcripts. But this hardly means that they are regulatory or otherwise functional. Indeed, if one carries the ENCODE definition to its logical extreme, 100% of the genome is functional—for all of it participates in at least one biochemical process—DNA replication.

That the ENCODE project would not adopt the most extreme biochemical definition is understandable—that definition would be useless. But the ENCODE definition is still grossly overinclusive from the standpoint of evolutionary biology. From that perspective, most estimates of the proportion of “functional” DNA are well under 80%. 4/
In short, evolutionary biologists reject "biochemical function" as a criterion for recognizing "junk" because not every bit of biochemical activity affects the reproductive fitness of organisms. (Neither does chemical activity per se show any influence on phenotypes that are related to the healthy functioning of those organisms.) To the evolutionary biologists, the term “junk DNA” means parts of the genome in which the particular DNA sequences (the order of the base pairs) do not have evolutionary significance. The Times article defines “junk DNA” differently, and vaguely, as “pieces of DNA that do nothing for us.” This is not the scientific definition. In fact, the earliest papers on “junk DNA” proposed that much of it might “do something” for us.

The “junk DNA war” (or rather the confusion about the meaning of the term “junk”) has spilled over into the legal realm. A brief that leading genetics and genomics researchers submitted to the U.S. Supreme Court to clarify the privacy implications of forensic DNA typing tried to address it. 5/ These researchers observed that
  • In genetics, “junk DNA” denotes sequences that lie outside of genes and that are not under detectable selective pressure: that such DNA exists is not in doubt.
  • “Junk” DNA sequences could be biologically useful or interesting yet not be useful for disease diagnosis or prediction.
  • ENCODE data do not reveal that anywhere near 80% of the genome contains medically relevant information.
  • The ENCODE findings indicate that the system that regulates gene expression is exquisitely complex, but they do little to change the status of “junk DNA” in general.
As far as I know, these conclusions have not been contradicted by new studies, but I have not conducted a recent literature review and would be grateful to hear of relevant papers that undermine these observations.

Notes
  1. Carl Zimmer, Is Most of Our DNA Garbage?, N.Y. Times Mag., Mar. 5, 2015 
  2. Re: Is Most of Our DNA Garbage?, N.Y. Times Sunday Mag., Mar. 20, 2015
  3. See Sean R. Eddy, The C-value Paradox, Junk DNA and ENCODE, 22 Current Biology R898 (2012)
  4. David H. Kaye, ENCODE’S “Functional Elements” and the CODIS Loci (Part II. Alice in Genomeland), Forensic Science, Statistics, and the Law, Sept. 18, 2012 (note omitted)
  5. Brief of Genetics, Genomics, and Forensic Science Researchers as Amici Curiae in Support of Neither Party, Maryland v. King, No. 12-204, Dec, 28, 2012, reprinted in part in Henry T. Greely & David H. Kaye, A Brief of Genetics, Genomics and Forensic Science Researchers in Maryland v. King, 53 Jurimetrics J. 43 (2013), available at http://ssrn.com/abstract=2403063http://ssrn.com/abstract=2403063. Disclosure statement: I prepared an initial draft of the brief and coordinated the revisions to it.

Thursday, March 5, 2015

The (Lack of) Meaning of the Supreme Court's Disposition of Raynor v. State

Yesterday, Popular Science reported that a “recent refusal by the Supreme Court means that involuntary DNA collection isn't unconstitutional.” This will come as a surprise to the Justices who voted to deny a writ of certiorari to the Maryland Court of Appeals in Raynor v. State, 99 A.3d 753 (Md. 2014).

Raynor is one of many cases in which courts have concluded that the Fourth Amendment prohibition against “unreasonable searches and seizures” does not apply to acquiring and testing naturally shed DNA. This particular case arose when, two years after a reported rape, the victim told police that she suspected Glenn Raynor had attacked her. Raynor agreed to come to a police station to answer questions. At the interview, he declined to provide a DNA sample, but after he left, police took swabs of the armrests of the chair in which had sat. The trial court denied his motion to suppress evidence of the incriminating match that followed, noting that “if he was so concerned about it, he should have worn a long sleeve shirt.” A conviction and a 100-year sentence of imprisonment followed.

According to the Popular Science article,
Raynor appealed the decision, saying the DNA evidence shouldn't have been used because it was collected without his consent. The appeal made it all the way up to the Supreme Court, which on Monday, the court announced [sic] that it would not hear the case. The Supreme Court did not comment on the denial—and to be fair, they get requests to hear a whole lot of cases every year and have to deny a majority of them—[but] their refusal to hear the case means they stand with the lower court’s majority opinion [which stated that]:
We hold that DNA testing of the 13 identifying junk loci within genetic material, not obtained by means of a physical intrusion into the person’s body, is no more a search for purposes of the Fourth Amendment, than is the testing of fingerprints, or the observation of any other identifying feature revealed to the public—visage, apparent age, body type, skin color.
In fact, the Supreme Court denies some 97% of the petitions it receives from private parties. Any first year law student knows that denying one of these 7,000 or so petitions does not mean that the Court “stand[s] with the lower court’s majority opinion.” It merely means that, for any number of possible reasons, four of the nine Justices did not vote to re-examine the case. In short, although police have been doing such testing time and again over the last twenty years or so, the U.S. Supreme Court has yet to approve — or disapprove — of the constitutionality of the practice.

References
Related posting

Tuesday, February 24, 2015

Genetic Determinism and Essentialism on the Electronic Frontier

The latest bit of what, in the scientific world, is discredited genetic determinism, comes from the Electronic Frontier Foundation (EFF). This is not the first time the EFF has strayed from electronics to genetics, where it seems inclined to overstate scientific findings. 1/ Now the organization wants the Supreme Court to decide whether it is an unreasonable search or seizure for police, without probable cause and a warrant, to acquire and analyze shed DNA for identifying features that might link a suspect to a crime. That is a perfectly reasonable request, although, in the unlikely event that the Court takes this bait, making the case for a Fourth Amendment violation will not be easy.

What is less reasonable, indeed, what many geneticists and bioethicists regard as ill-advised, is to portray DNA as a map of “who we are, where we come from and who we will be.” 2/ My DNA is not who I am. It determines some things about me — my blood type, for example — but not my occupation, my interests, my skills, my criminal record, or my political affiliation. Yet, rather than simply point out that people have legitimate reasons to want to maintain the confidentiality of certain traits or risks that DNA analysis could reveal — such as an inherited form of Alzhiemer’s Disease — the EFF is concerned that “[r]esearchers have theorized DNA may also determine race, intelligence, criminality, sexual orientation, and even political ideology.” 3/

Of course, researchers have “theorized” almost everything at one time or another. And the prospect that police will collect DNA from a suspect surreptitiously to find out if he is a liberal Democrat or a conservative Republican seems a tad silly. Still, I was curious: Is there really a theory of how genes determine political ideology?

I turned to the news article in a 2012 issue of Nature cited by the EFF. 4/ Nothing in the article gives a theory of genetic determinism for political ideology. The article refers to twin studies that imply genetics plays some role in political behavior. There are some reports of candidate genes from studies that have “yet to be independently replicated.” 5/

As for a theory of how unknown genes might, to some degree, in some settings, influence political ideology, the theory is that some genes affect general attitudes or emotional reactions that could relate in some manner to political ideology. For example,
US conservatives may not seem to have much in common with Iraqi or Italian conservatives, but many political psychologists agree that political ideology can be narrowed down to one basic personality trait: openness to change. Liberals tend to be more accepting of social change than conservatives. ...

Theoretically, a person who is open to change might be more likely to favour gay marriage, immigration and other policies that alter society and are traditionally linked to liberal politics in the United States; personalities leaning towards order and the status quo might support a strong military force to protect a country, policies that clamp down on immigration and bans on same-sex marriage. 6/
These remarks are not a basis for a true friend of the Court to imply that political ideology might be a genetically determined phenotype. 7/

Notes
  1. See David H. Kaye, Dear Judges: A Letter from the Electronic Frontier Foundation to the Ninth Circuit, Forensic Science, Statistics and the Law, Sept. 20, 2012.
  2. Brief of Amicus Curiae Electronic Frontier Foundation in Support of Petitioner on Petition for a Writ of Certiorari, Raynor v. Maryland, No. 14-885, Feb. 18, 2015, at 2.
  3. Id. (note omitted).
  4. Lizzie Buchen, Biology and Ideology: The Anatomy of Politics, 490 Nature 466 (2012).
  5. Id. at 466.
  6. Id. at 468.
  7. For a critical discussion of factual errors and distortions in Supreme Court amicus briefs generally, see Allison Orr Larsen, The Trouble with Amicus Facts, 100 Va. L. Rev. 1757 (2014).

Friday, February 20, 2015

Buza Reloaded: California Supreme Court Grants Review

Yesterday the California Supreme Court granted review in People v. Buza, No. A125542 (Cal. Ct. App., 1st Dist., Dec. 3, 2014), and ordered the Court of Appeal opinion "depublished." A depublication order "is not an expression of the court's opinion of the correctness of the result of the decision or of any law stated in the opinion." Cal. Rules of Court, Rule 8.1125(d) (2015). However, "an opinion of a California Court of Appeal ... that is not ... ordered published must not be cited or relied on by a court or a party in any other action" in California. Rule 8.1115(a).

The California Department of Justice issued a information bulletin advising all state law enforcement agencies that
By operation of state law, the Supreme Court’s order granting review removes the Court of Appeal’s opinion as published authority and prevents citation or reliance on that decision in any other action. As a result of the California Supreme Court’s grant of review of this decision, there is now no state precedent that precludes collection of DNA database samples from adult felony arrestees pursuant to Penal Code section 296.

Penal Code sections 296(a)(2) and 296.1(a) therefore are in full effect and mandate the collection of DNA database samples from all adults arrested for a felony or wobbler offense. All authorized arrestee samples that have been or will be received by the California Department of Justice DNA Data Bank program will be analyzed and uploaded to CODIS.

Closely related postings

Sunday, February 15, 2015

"Remarkably Accurate": The Miami-Dade Police Study of Latent Fingerprint Identification (Pt. 2)

A week ago, I noted the Justice Department’s view that a “study of ... latent print examiners ... found that examiners make extremely few errors. Even when examiners did not get an independent second opinion about the decisions, they were remarkably accurate.” 1/ But just how accurate were they?

The police who conducted the study “[p]resented the data to a professor from the Department of Statistics at Florida International University” (p. 39), and this “independent statistician performed a statistical analysis from the data generated” (p. 45). The first table in the report (Table 4, p. 53) contains the following data (in slightly different form):

Table 1. Classifications of Pairs
Examiner's
Statement
Nonmates (N) Mates (M)
953 235
+ 42 2547
? 403 446

Here, “+” stands for a positive opinion of identity between a pair of prints (same source), “–” denotes a negative opinion (an exclusion), and “?” indicates a refusal to make either judgment (an inconclusive) even though the examiner initially deemed the prints sufficient for comparison.

What do the numbers in Table 1 mean? As noted in my previous posting, they pertain to the judgments of 109 examiners with regard to various pairings of 80 latent prints with originating friction ridge skin (mates) and nonoriginating skin (nonmates). A total of 3,138 pairs were mates; of these, the examiners reached a positive or negative conclusion in 2,692 instances. Another 1,398 were nonmates; of these, the examiners reached a conclusion in 995 instances. Given that examiners were presented with mates and that they reached a conclusion of some sort, the proportion of matches declared was P(+|M & not-?) = 2,457/2,692 = 91.3%. These were correct matches. For the pairings in which the examiners reached a conclusion, they declared nonmates to match in P(+|N & not-?) = 42/995 = 4.2% of the pairs. These were false positives. With respect to all the comparisons (including the ones that they found to be inconclusive), the true positive rate was P(+|M) = 2,457/3,138 = 78.3%, and the false positive rate was P(+|N) = 42/1,398 = 3.0%. Similar reasoning applies to the exclusions. Altogether, we can write:

Table 2. Conditional Error Rates

Excluding inconclusives Including inconclusives
False + P(+ | N & not-?)
4.2%
P(+ | N)
3.0%
False – P(– | M & not-?)
8.7%
P(– | M)
7.5%


These error rates, which are clearly reported in the study, do not strike me as "remarkably small"—especially considering that they include the full spectrum of pairs—easy as well as difficult comparisons. Of course, they do not include blind verification of the conclusions, a matter addressed in another part of the study.

The authors report more reassuring values for “Positive Predictive Value” (PPV) and “Negative Predictive Value (NPV).” These were 98.3% and 92.4%, respectively. But these quantities depend on the proportions of pairs that are mates (69%) and nonmates (31%) in the test pairs. The prevalence of mates in casework—or the “prior probability” in a particular case—might be quite different. 2/

A better statistic for thinking about the probative value of an examiner’s conclusion is the likelihood ratio (LR). Are matches declared more frequently when examiners encounter mated pairs than nonmates? How much more frequent are these correct classifications? Are declared exclusions more frequent when examiners encounter nonmates than mates? How much more frequent are these correct classifications?

The LR answers these questions. For declared matches, the LR is P(+|M) / P(+|N) = 0.783 / 0.030 = 26. For declared exclusions, it is P(–|N) / P(–|M) = 9. 3/ These values support the claim that, on average, examiners can distinguish paired mates from paired nonmates. If all the examiners were flipping fair coins to decide, the LRs would be expected to be 1. The examiners did much better than that.

Nevertheless, claims of overwhelming confidence across the board do not seem to be justified. If examiners were presented with equal numbers of mates and nonmates, one would expect that a declared match would be a correct match in P(M|+) = 26/27 = 96% of the cases in which a match is declared. 4/ Likewise, a declared exclusion would a correct classification in P(N|–) = 9/10 = 90% of the instances in which an exclusion is declared. The PPV and PNV in the Miami-Dade study are a little bit higher because the prevalence of mates was 69% instead of 50%, and the examiners were cautious — they were less likely to err when making positive identifications than negative ones.

Suppose, however, that in a case of average difficulty, an average examiner declared a match when the defendant had strong evidence that he never had been in the room where the fingerprints were found. Let us say that a judge or juror, on the basis of the non-fingerprint evidence in the case, would assign a probability of 1% rather than 50% or 69% to the hypothesis of that the defendant is the source of the latent print. The examiner, properly blinded to this evidence, would not know of this small prior probability. An LR of 26 would raise the prior probability from 1% to 26%. Informing the judge or juror of the reported PPV of 98.3% from the study without explaining that it does not imply a “predictive value” of 98.3% in this case would be very dangerous. It would lead the factfinder to regard the examiner’s conclusion as far more powerful than it actually is.

Notes

  1. David H. Kaye, "Remarkably Accurate": The Miami-Dade Police Study of Latent Fingerprint Identification (Pt. 1), Forensic Science, Statistics, and the Law,  Feb. 8, 2015
  2. In addition, the NPV has been adjusted upward from 80% “[i]n [that] consideration was given to the number of standards presented to the participant.” P. 53.
  3. Removing nondeclarations of matches or exclusions (inconclusives) from the denominators of the LRs does not change the ratios very much. They become 22 and 11, respectively.
  4. This result follows immediately from Bayes' rule with a prevalence of P(M) = P(N) = 1/2, since P(M|+) = P(+|M) P(M) / [P(+|M) P(M) + P(+|NM) P(NM)] = P(+|M) / [P(+|M) + P(+|NM)] = LR / (LR + 1) = 26/27.