
WHEN tests were developed to identify criminals by ‘typing’ or profiling
their DNA, many lawyers and forensic scientists heralded them as ‘a prosecutor’s
dream’ and ‘the greatest advance in crime-fighting technology since fingerprints’.
The breakthrough that led to the creation of DNA ‘fingerprints’ was reported
in 1985 by Alec Jeffreys of the University of Leicester; two years later
Britain and the US introduced the tests into their legal systems. Recently,
however, a cloud of doubt has fallen on the technology that appeared to
be a panacea for fighting crime. There is a growing suspicion that, in the
US at least, the new DNA tests were rushed to court before they were ready.
Enthusiasm for DNA fingerprints to identify criminals was such that,
in the US alone, three commercial laboratories were offering testing services
by late 1987. The FBI then instigated a crash programme to develop its own
procedure for DNA typing which it has been offering to police departments
and law enforcement agencies since last year. Local crime laboratories in
several states are also setting up their own facilities. To date, DNA tests
have helped to win convictions in more than 100 criminal cases. At first,
the use of the test went virtually unchallenged in the courts. Several attorneys
reported that they could not find a single scientist who questioned the
accuracy of the evidence incriminating their clients.
Last year, however, defence attorneys started to find that scientists
were questioning the claims of the laboratories concerning the accuracy
and reliability of their tests. The criticisms are not so much with the
underlying theory, as the lack of standards for carrying out the test and
interpreting the results. The scientists also claim that proficiency studies
– to demonstrate that the DNA tests are reliable on forensic samples – have
been woefully inadequate. Not enough testing has been done and the tests
have not involved realistic forensic samples.
Advertisement
The underlying scientific dispute has surfaced in courtrooms in a number
of states during hearings to determine whether the new DNA tests are ‘accepted
in the scientific community’. Under a rule first established in 1923 in
the case of United States v. Frye, a new scientific test or procedure may
be presented in a jury trial only if it has ‘gained general acceptance in
the particular field in which it belongs’. The rule is based on the assumption
that lay jurors are not able to evaluate scientific procedures and will
tend to accept anything ‘scientific’ as infallible. The legal system itself,
goes the argument, should try to assure that scientific evidence is reliable
before allowing the evidence to be presented to the jury. The best way to
do this, the law assumes, is to determine in a pre-trial hearing whether
knowledgeable scientists accept the technique as reliable. Most American
states follow the Frye rule, and in those states a ‘Frye hearing’ must be
held if the defence challenges a new scientific test, such as DNA typing.
Last August, in the widely reported murder case of People v. Castro
in New York, Judge Gerald Sheindlin ruled that a test performed by Lifecodes,
the largest of the commercial forensic DNA laboratories in the US, could
not be used in evidence against the defendant. The ruling came after a Frye
hearing in which the defence attorneys, Peter Neufeld and Barry Scheck,
presented testimonies by a number of distinguished scientists led by Eric
Lander, a geneticist at the Massachusetts Institute of Technology in Cambridge
near Boston. Lander declared the procedures for interpreting results at
Lifecodes ‘so far below reasonable scientific practice in molecular biology
as to be appalling’. Two experts who had testified for the prosecution,
Richard Roberts, a molecular biologist at Cold Spring Harbor Laboratory,
New York, and Carl Dobkin, a research scientist in molecular biology for
New York State, were sufficiently swayed by Lander’s evidence to withdraw
their endorsement of Lifecodes. Roberts and Dobkin then joined Lander in
signing a statement which declared the laboratory’s results ‘not scientifically
reliable enough’ to support its conclusion that blood found on the defendant’s
watch ‘matched’ the blood of a woman whom the defendant was accused of stabbing
to death.
Widespread publicity about the Castro case prompted other defence lawyers
to arrange for independent experts to scrutinise the results of DNA tests
on their clients. In many instances, the independent experts found that
the conclusions of the DNA laboratories were not adequately supported. In
the face of serious challenges by defence attorneys and their experts, prosecutors
in a number of states including California, Pennsylvania, Massachusetts,
Arizona and Louisiana, chose to withdraw the results of DNA tests that they
had offered in evidence.
For example, the FBI found its DNA test being challenged seriously for
the first time in September by an Arizona court when Alvie Kiles stood trial
for three murders. The prosecution, after presenting testimony by several
witnesses who supported the test, suddenly announced it would withdraw the
FBI’s DNA evidence. This unexpected action apparently resulted from a court
order that the FBI disclose the results of proficiency studies conducted
in its laboratory. The judge announced that he would rule the test inadmissible
unless the FBI complied with the order; the next day the evidence was withdrawn.
Such decisions, however, lacked any value as legal precedent. Until
a higher court in a state, such as the supreme court, issues a ruling on
the admissibility of DNA evidence, a Frye hearing is necessary every time
the defence challenges the DNA test. Such a precedent was set in November,
when the Supreme Court in Minnesota ruled as inadmissible the tests by Cellmark
Diagnostics, a subsidiary of ICI and the second largest commercial forensic
DNA laboratory in the US. Although the court found that DNA typing is accepted
in the scientific community, it held that Cellmark’s test was not ready
to be presented in evidence. In the opinion of the Supreme Court, Cellmark
had failed to meet minimal standards for the scientific validation of its
laboratory protocol. The company had also refused to disclose information
needed to evaluate the accuracy of its test, the court said, and it had
failed to publish data on its procedures in peer-reviewed scientific journals.
The decision on Cellmark’s test is binding on Minnesota courts, and no doubt
will be carefully studied in other states.
To understand the arguments in the dispute as to the accuracy of the
method, it is necessary to understand the theory behind DNA fingerprinting.
DNA is a long, chain-like structure composed of pairs of molecules called
bases. There are four types of bases and, in human DNA, more than a billion
base pairs. The sequence of base pairs constitutes the genetic code responsible
for inheritance: no two individuals (except identical twins) have exactly
the same number or sequence of base pairs. The test is based on the premise
that, as an individual’s DNA is the same regardless of whether it comes
from hair, blood, sperm, muscle or nerve tissue, a comparison of the DNA
from any tissue from two individuals should identify many differences. But
there would also be many similarities, however. According to some estimates,
more than 90 per cent of the DNA in humans is the same from one individual
to the next. Indeed, some DNA is common to both animals and humans; in the
case of chimpanzees, well over half is the same. The trick with DNA typing
is to examine those parts of the DNA molecule where there tend to be differences
rather than the far more numerous places which are the same.
The most common approach focuses on sections of DNA where certain short
sequences of base pairs are repeated, over and over. If the chain of base
pairs in a DNA molecule were a phonograph record, these areas would be points
where the needle became temporarily stuck in a groove and repeated the same
notes a number of times before playing the rest of the tune. Because the
number of repeats tends to vary among individuals, the areas are known as
‘variable number tandem repeats’ or VNTRs.
The most widely used DNA typing tests rely on a technique called restriction
fragment length polymorphism (RFLP) analysis to measure the length of fragments
of DNA containing VNTRs. In this, the laboratory extracts DNA from a sample,
cuts the long, chain-like molecules into fragments which are sorted by length.
It identifies those fragments which contain a particular VNTR and measures
them – the number of repeats in the VNTR will dictate the length.
If the analyst compares two fragments containing a given VNTR and finds
a difference in length, then it is clear that the samples could not have
come from the same person. If the length of the fragments containing the
VNTR are the same, then the samples could have come from the same person,
but they could also have come from another individual who just happens to
have a VNTR of the same length. If the lab compares several fragments containing
different VNTRs, however, the likelihood of a coincidental match on all
of them is decreasingly small.
Although DNA typing is simple in theory, this approach is complicated
in practice. The tests require several distinct steps and they can take
up to six weeks to perform .
DNA typing is relatively new for criminal identification, but RFLP analysis
is widely used in research and in medical diagnostics. So those who favour
the results of forensic DNA typing being used as evidence in court – typically
prosecutors – argue that the new tests are simply a new application of a
well-established scientific procedure. Those who oppose the technique’s
use in court – typically defence attorneys – argue that DNA typing is a
more difficult, exacting procedure than other applications of RFLP analysis
and that the results are more difficult to interpret.
First, say the opponents, the probes used for criminal identification
locate DNA fragments that are ‘hypervariable’: the number of times a sequence
of base pairs is repeated in a VNTR can vary from one to several hundred.
Each fragment of a given length is a specific genetic variant or allele,
but it may differ in length from other fragments or alleles by such a small
increment as to be almost indistinguishable. Precise measurement is critical;
minor differences in the position of bands on a DNA print may mean the difference
between conviction and acquittal. By contrast, the probes used to diagnose
such diseases as Huntington’s disease or cystic fibrosis identify fragments
that are less variable; typically only four alleles need to be distinguished.
These alleles produce such highly distinctive patterns that the length of
the underlying fragments does not need to be measured.
Secondly, diagnostic and research labs generally carry out RFLP analysis
on DNA extracted from fresh blood and tissue samples under carefully controlled
conditions. In these cases, the scientist knows precisely what is in a given
sample. This is not the case in forensic work, where samples may contain
virtually anything. A blood stain on dirty fabric found at the scene of
the crime may contain a variety of chemical and biological contaminants,
such as bacteria, detergents, dyes and dirt as well as DNA from other humans
or even from animals. The sample may also be degraded by age, exposure to
the environment or putrefaction. Forensic laboratories attempt to purify
their samples, but DNA purification can be a tricky business, particularly
when the samples are tiny and contain unknown contaminants.
Where the errors creep in
One result of contamination is that, while DNA prints of different individuals
may look quite similar, prints from the same individual often look different.
The popular perception of DNA prints is that they resemble supermarket bar
codes, but those produced by the DNA labs rarely have the crisp, clean appearance
of their supermarket counterparts. The bands are often smudgy and smeared,
making it difficult to tell where one bands starts and another ends. Comparison
of two prints is often further complicated by aberrations that can arise
at several stages in the RFLP analysis because of poor samples or sloppy
laboratory practice.
During ‘restriction digestion’, for example, the DNA must be cut precisely
the same way each time, because the test depends on a comparison of the
length of the resulting fragments. But restriction digestion is a temperamental
process. Careless laboratory procedure or the presence of contaminants in
a sample can cause too few or too many cuts to be made, spuriously altering
the length of fragments to be compared. The result is a DNA print with too
many bands or with bands in the wrong places. Alternatively, contaminants
or carelessness during electrophoresis can cause a sample to move more quickly
or slowly through the gel than the sample with which it is being compared.
The result is an aberration known as band shift, in which samples from the
same individual produce DNA prints with similar patterns but with the bands
out of alignment. One print appears to be shifted slightly up or down from
the other. Finally, extra bands can appear at the hybridisation stage if
the probes are contaminated, if they lock onto nonhuman DNA or if the probes
hybridise to the wrong fragments. In other instances, the probes may fail
to lock onto human DNA because the sample is contaminated or degraded, with
the result that a band which should appear is absent. Band shifts, extra
bands and missing bands have been persistent problems in forensic casework.
Proponents of DNA typing acknowledge these problems, but point out that
the aberrations are far more likely to make matching DNA prints look different
than to make different DNA prints look the same. They may exonerate a guilty
individual, but they are unlikely to incriminate an innocent person. The
problem with this argument is that the DNA laboratories, knowing that such
problems can and do arise, have tried to make allowances to avoid too many
guilty suspects being falsely exonerated. Experienced analysts sometimes
claim to be able to distinguish true and aberrant bands on the basis of
their appearance. Such practices add a significant element of subjectivity
and guesswork to the interpretation of the results.
For example, in a multiple-rape case in Buffalo, New York, Lifecodes
matched a blood sample from the defendant to semen stains taken from five
different victims. After initially declaring a match, the company retested
the defendant and found a different pattern. Lifecodes attributed the discrepancy
to an error in the first print: a band that should have appeared was missing
for unknown reasons, and an ‘extra’ band appeared at a different location
due to a problem in restriction digestion. Did this invalidate the previous
matches? No. The standards for declaring a match were such that Lifecodes
could conclude that the second DNA print, although different from the first,
also matched the five semen stains.
In another case, a murder trial in Atlanta, a representative for Lifecodes
testified that two DNA prints were a ‘match’ even though none of the bands
was in alignment. His reasoning: ‘There is, however, a consistent non-alignment
of the bands throughout the test, telling us there’s a match.’ While this
reasoning has its appeal, and may prevent false exclusions due to band shift,
it expands the range of DNA prints that can be called a match. As a result,
it is less likely that the test will correctly exclude the innocent, undermining
its power of identification to a degree that is difficult to estimate.
Lawyers’ uncertainty over the proper interpretation of DNA prints might
not be so great if the tests were repeated several times to assure the reliability
of a particular finding. However, forensic DNA typing, unlike other applications
of RFLP analysis, relies on a single, unrepeated test.
How likely is it that laboratory error or misinterpretation will produce
a mistaken result? The sad truth is that we do not really know. In the only
meaningful trial conducted to date by an independent organisation, all three
commercial DNA laboratories made errors. Asked to test approximately 50
unknown blood and semen samples, both Cellmark and a second laboratory,
Forensic Science Associates, had false positives. That is, they declared
a match where the pair of samples under comparison came from different people.
Lifecodes had no false positives, but was unwilling to make a decision on
14 of the samples and, in a follow-up study, twice failed to detect that
mixed stains contained the DNA of two individuals. This sort of error rate
hardly justifies the laboratories’ claim of ‘exquisite’ accuracy.
Recently, scientists have also begun questioning the techniques that
the commercial laboratories use to compute the statistics they present in
connection with DNA evidence. At issue is the laboratories’ claim that only
one person in millions or billions is likely to have a given DNA print.
To arrive at these figures, the companies have analysed blood samples from
several hundred people and recorded the position of the bands in each person’s
DNA print in a ‘population database’. They then calculate the probable frequency
of any one band in the population as a whole by determining the percentage
of bands in the database in a corresponding position. To estimate the frequency
at which a given DNA print will occur, they multiply together the frequencies
of the bands produced by each probe.
ÐÓ°ÉÔ´´s have attacked the statistics on two levels. The first is
the assessment of the frequency of the various bands in the DNA print. As
we have said, DNA prints of the same individual often look a bit different.
To take this variability into account, and thereby avoid falsely exonerating
the guilty, the laboratories allow a certain amount of latitude when calling
two bands a match. Experts testifying for the defence in a recent hearing
in California noted that Cellmark’s estimate of frequency was several orders
of magnitude lower than if it had used the method advocated by the FBI laboratories.
Another expert offered similar testimony with respect to Lifecodes in the
People v. Castro case in the Bronx.
The second factor that scientists attack is the assumption that people
mate at random (with respect to the alleles used in DNA typing) and, therefore,
that the various bands are distributed more or less at random across a homogenous
population. The chief concern is that the population is not homogenous in
the US, but instead ‘structured’ so that certain sets of alleles are more
common in some population subgroups than in others. If this is the case,
a particular DNA print may be far more common in some subgroups than others,
and the DNA laboratories may incorrectly estimate its frequency by sampling
from the wrong group when constructing their databases.
One index of the extent of structuring in the population is the number
of homozygotes: that is, the number of individuals who have one rather than
two bands for a given probe because they received that same allele from
each parent . If people who are genetically similar tend to mate with each
other, the number of homozygotes increases. Several population geneticists
have recently examined the databases of Lifecodes and Cellmark. They conclude
that these databases contain many times the number of homozygotes that would
be expected in a population mating at random. Another indication of population
structuring was found in the FBI’s Hispanic database, which contained data
from Miami and Houston. There were striking differences between the two
cities in the frequency of many of the alleles.
These are not insurmountable problems, and many are now being addressed.
For example, the Congressional Office of Technology Assessment and the National
Academy of Sciences have appointed panels of experts to recommend standards
for DNA testing. Both panels are due to report before the end of the year.
The DNA laboratories themselves are working to complete validation studies
and to develop new control procedures to solve the problems with restriction
enzymes, band shift, degradation and other causes of laboratory error. The
most urgent need is for rigorous proficiency testing to determine just how
often errors occur in DNA typing. At present, however, there is no government
or scientific agency whose responsibility it is to set up proficiency tests;
nor is there much incentive for laboratories to submit to such testing.
The necessary testing is likely to be done only if new legislation is passed,
or if supreme courts insist that it be done before DNA test results are
admissable as evidence in court.
Because DNA tests came into routine use in the US before these steps
were completed, however, the laboratory procedures adopted are not stringent
enough to prevent the possibility of errors – for example, mislabelling,
contamination of samples and contamination of probes. Interpretive standards
either do not exist or are not consistently followed. The fact that DNA
testing was introduced prematurely is, at least in part, a result of genuine
enthusiasm for a technique. The laboratories, in the rush to meet the market
for forensic tests, may have made too many simplifying assumptions and cut
too many corners. The outcome, though, was that important scientific issues
were left to be worked out in courtroom hearings with justice and the fate
of criminal defendants hanging in the balance.
* * *
From sample to DNA profile in the forensic laboratory
THE first step in DNA testing is for the laboratory, by various chemical
procedures, to extract DNA from the samples. It then adds a restriction
enzyme which ‘cuts’ or digests the DNA at defined points to break the long
chain: the DNA now resembles a pile of twisted ribbon cut into fragments
of various lengths.
The fragments are sorted by length using a process known as electrophoresis.
The analyst places the cut DNA in a slab of agarose gel (made from agar,
found in kelp) and runs an electric current through the gel with the negative
electrode at the end near the samples. The DNA fragments move toward the
positive electrode, but the shorter ones move more rapidly through the gel
than the longer ones.
To sort those fragments containing a particular VNTR from the other
fragments in the gel, the analyst transfers the contents of the gel to a
nylon membrane using a method known as Southern blotting (in honour of Edward
Southern, a biologist from the University of Edinburgh, who developed it).
The membrane is placed on the top of the gel and covered with absorbent
paper towels. Capillary action draws the DNA fragments directly upwards
so that they stick to the membrane with their relative positions unchanged.
The next step is to ‘hybridise’ the DNA. The membrane is immersed in
a solution of radioactive probes which are capable of finding and binding,
or hybridising, them selves to the fragments of DNA that contain the VNTR
of interest. A probe binds to a DNA fragment only if it finds a specific
sequence of base pairs; its structure is complementary to the sequence it
is designed to locate.
The final step is to make the VNTR fragments visible by placing the
membrane in contact with X-ray film. The probes contain a radioactive molecule,
so wherever a probe has bound to a DNA fragment it emits radiation which
exposes the film, producing a band. If two samples have VNTRs of the same
length, the bands will appear in the same relative position on the films.
The resulting X-ray films are called autoradiographs (autorads or rads for
short) and the patterns of bands are known as ‘DNA prints’.
Most of the probes used in DNA typing produce either one or two bands
in a DNA print. Humans have two copies of most chromosomes; one from the
mother, the other from the father. Consequently, people have two copies
of a given VNTR. If the two copies are of different lengths, two bands will
appear on the autorad; if they are the same length, one band will appear.
Individuals who have two bands are called heterozygotes; those with one
band are homozygotes.
Forensic DNA labs use three or four probes, which produce DNA prints
consisting of up to six or eight bands. To determine whether two DNA prints
match, the analyst checks the autorad to see whether the bands have the
same pattern and whether they are aligned. But visual inspections can be
deceptive, so the analyst confirms each apparent match by calculating the
length of the DNA fragment denoted by each band. This involves measuring
the position of each band relative to other bands produced by a ‘marker’
DNA on the same gel and which contains DNA fragments of known size. A DNA
profile consists of a set of six to eight numbers, each indicating the position
of a band in the DNA print.
William C. Thompson is a lawyer and social scientist and Simon Ford
a molecular biologist at the University of California, Irvine.