鈥榃E NEED very much a name to describe a cultivator of science in general. I should incline to call him a 杏吧原创.鈥 This is how the Oxford English Dictionary explains the way in which, and when (1840), William Whewell invented a word that is now in common use. Whewell was rather special when it came to words: his name is attached to quotations supporting 607 words in the dictionary, rather more than one might expect from the average cultivator of science. All this is revealed by the dictionary, whose second edition was published this spring and launched to something like rapture from the literary world. 鈥楲ike being present at the revelation of the second edition of the Bible,鈥 said Daniel Boorstin, former Librarian of the US Congress, with the sort of artistic flourish that tends to leave science shuffling tongue-tied in the corner.
杏吧原创s and words have never had an easy relationship. It sometimes seems like a one-sided tussle as scientists struggle to find the words to express themselves in writing. It may often look as though language has defeated science, but as the dictionary shows, eventually it is language that must accommodate science. What makes the OED so interesting to scientists is that English has become the language of science. The language has not always held this position: German in particular, and French also, have challenged it in the past. When the Pasteur Institute, in Paris, decided in March this year to publish its Annales in English, it confirmed how established the language has become. (The move caused a political storm in France. Axel Kahn, director of research at INSERM, the state funding agency for medical research, complained bitterly in Le Monde of a cultural and economic imperialism that would deny French scientists the right to think and write in their native language. But even he accepts that English is the 鈥榩rincipal scientific language鈥.) Among dictionaries, the OED has an untouchable reputation. In sheer size, scope and scholarship it will never be equalled. In 22 000 pages, it defines more than half a million words, many of them obscure or obsolete, but all with one thing in common: they form part of the English language. Every word, or meaning of a word, is supported by quotations from the first recorded use, backed up by one or more further quotations that amplify the description. The great bulk of the work, the first edition covering some 400 000 words and phrases, began to appear in 1884. It had taken the five years from 1879 to produce the first volume, 鈥楢-ANT鈥. The last volume was published in 1928. One supplement appeared in 1933. A new supplement, absorbing the first, was published in four volumes from 1972 to 1986.
The second edition, or OED2, takes in all the previous material, adding a further 5000 words and meanings. This means that the dictionary is also a history of science, recording the birth and growth of almost every word we use. The entry for 鈥榓lgebra鈥, for example, will tell you among other things that it comes from an Arabic word for reuniting broken parts and is closely related to bone-setting; that as 鈥榯he science of redintegration and equation鈥 it entered the Italian language in 1202; that it appeared in English as a word for 鈥榯he surgical treatment of fractures鈥, and in its current sense in 1551.
Advertisement
Award-winning programs
The OED2 is also a dictionary from science. The sheer labour involved in producing a new edition meant that computers, in various forms, were indispensable to the project. The entire content of the first edition and its supplements was first captured electronically, a total of around 350 million characters. Specially written software then integrated the entries from the different texts into one. The development of the software is recognised as a massive achievement. James Howes, chief computer systems designer, said the first question was to decide if it could even be done 鈥榳ithin anything like a reasonable time鈥. But the computerisation was completed on time and to budget, and it won the Applications Award from the British Computer Society in 1987.
The team from Oxford University Press worked with the British arm of IBM and the University of Waterloo in Canada, as well as receiving a grant from the Department of Trade and Industry. IBM itself now uses some of the software developed for the project. For the University of Waterloo, a new university that was built with computers for staff and students in mind and which is heavily involved in work on computer science, it was a chance to apply research to the world of work. Both IBM and the University of Waterloo seem delighted with the result. Significantly, both focus on the relationship between science and the arts. 鈥楾he greatest untapped asset of computer technology will be in the arts and the humanities,鈥 said Geoffrey Robinson from IBM at the OED2鈥檚 launch in London at the end of March. Douglas Wright, from the University of Waterloo, spoke of a dream that computers may one day prove more useful with words than they are now.
In the words of the lexicographers, the OED is 鈥榙escriptive鈥 rather than 鈥榩rescriptive鈥: it does not lay down how a word is to be used, it simply describes how the word has been used. In that sense, it treats English as a democratic language in which the only criterion for acceptance is whether a word, or a meaning, has come into popular use. So anyone can give birth to a new word. But it will not find its way into the dictionary unless English-speaking people adopt it.
Science has done more than enable the new dictionary to appear: it has also provided much of the new content. Of the 5000 new entries around 1200 are words of science or technology. Politics managed a scant 175 new words, while economics could muster only 95. That in itself should not be surprising, least of all to scientists. What might surprise more is that the dictionary acquired its first scientifically qualified lexicographer, Alan Hughes, only in 1968. Hughes read physics at Oxford, going on to work for Alcan as a physicist. But his eye was caught by Oxford University Press鈥檚 advertisement looking for a science graduate, and he was hooked. Now Hughes has four other scientists working with him, sitting in judgment on the words that science attempts to bring into the English language.
Immortality beckons for the progenitor of a word that the editors decide to accept, since the dictionary lists the author and work where the word first appeared. Hughes gives as an example HIV, the virus that causes AIDS, which just missed the OED2 (even with computerised production, dictionaries take time to put together). When the entry was initially prepared for the next edition, the first reference was to have been a paper by J. Coffin et al in Nature on 1 May 1986: 鈥榃e propose that the AIDS retroviruses, be officially designated as the human immunodeficiency viruses, to be known in abbreviated form as HIV.鈥 But after intensive research, pride of place 鈥 unless anyone comes up with an even earlier reference 鈥 will now go to Capital Gay, which wrote on 11 April: 鈥楢n international committee on viral names has been looking into the problem, and was rumoured to have agreed on 鈥榟uman immune deficiency virus鈥 (HIDV or HIV).鈥 The entry will remain even if the word later drops out of use, something which makes the dictionary a treasure trove for historians, both of language and of science. Most reference works will boast that they have been, they are 鈥榓ll new鈥 or 鈥榗ompletely revised鈥, but the OED2 stands as a record of what has been as well as what is.
From his office in St Giles鈥 in the heart of Oxford, Hughes coordinates a team of freelance readers who sift through a selection of well-known scientific books and journals, such as the British Medical Journal. Three publications are read in the office, word by word: Nature, New 杏吧原创 and Scientific American. (Quotations from New 杏吧原创 are attached to more than 1800 entries in the new OED.) That still leaves a large number of journals unscanned, and Hughes is the first to admit there is a 鈥榙egree of randomness鈥 about the reading programme. But it seems to work.
When Hughes (or one of his readers) spots a new word, phrase or meaning, he highlights it with a marker and passes it to a typist to type up a 鈥榪uotation slip鈥. This card is then filed, 鈥榝or eternity鈥 as Hughes confidently puts it. Sorters regularly go through the quotation cards, checking to see if they represent genuinely new words or meanings, in which case they can be promoted to become 鈥榖lue cards鈥. These are words that stand a good chance of making it into the next edition of the Shorter Oxford Dictionary or of the full dictionary. Reaching for a card index, Hughes reeled off 鈥榤acroassembler鈥, 鈥榤acroradical鈥 and 鈥楳RV鈥 (for multiple re-entry vehicle) as likely candidates. Each year his team writes some 15 000 quotation slips.
The minimum qualifications for entry are usage by more than one person, and at least three quotations. Just one quotation will guarantee only eternal oblivion in the files. So 鈥榩inastric acid鈥, known only from one book on lichens, and 鈥榩inocamphone鈥, which appeared only in Chambers Technical Dictionary in 1958, will not appear in the dictionary (unless some kind soul adopts them in print). 鈥業鈥檓 not saying these words don鈥檛 exist, just that their currency is very low,鈥 Hughes says. They share their fate with 鈥榩ink random vibration鈥, known solely from a glossary published by the British Standards Institution in 1976: 鈥楾hree-word phrases have a hard time,鈥 says Hughes unsentimentally.
Even so, it is for the editors to decide if and when to include a word. 鈥楢norexic鈥, for example, managed only to enter the OED2 even though its first recorded reference is given as in 1907. The baud, used for measuring the speed at which electronic data are transmitted, first appeared in print in 1929, five years before 鈥榖enzodiazepine鈥. Both have now finally made it into the dictionary. And 鈥榣arvivorous鈥 failed once again to gain entry despite an ancestry dating back a hundred years to 1889 (friends of the word will be relieved to know that it is set to appear in the next edition).
A first time for every word
One reason for delay 鈥 though not on the scale of a century 鈥 is that the OED does more than describe, and the editors do not rely solely on their own quotation slips. Once they have decided a word is suitable, they have to research it to ensure that the first reference really is the first. The list of words awaiting research includes some that recent events have made familiar enough for the television news, such as 鈥楥FC鈥 and 鈥榞reenhouse gas鈥, as well as the kind of 鈥榗ard鈥 you would slot into a computer and the 鈥楰/T boundary鈥 (which divides the Cretaceous and the Tertiary periods). Oceanographers might like to know that 鈥楨l Nino鈥, the abnormal warming in the Pacific, had been researched ready for the next dictionary. Hughes believes that it was first described in English in 1918, in the Geographical Review: 鈥業t is not uncommon for a current from the north 鈥 locally known as El Nino from its frequency during the Christmas season 鈥 to prevail in the region of Tumbes at 3.5Degree S鈥
What of the future? The first edition, the one that was finished in 1928, is now out of print but is available on CD-ROM. The publishers claim that with it you can generate, say, lists of words labelled as 鈥榤edical鈥 supported by quotations between 1700 and 1715. CD-ROM also turns the 1 827 306 separate quotations into a dictionary of quotations. The second edition should follow the first onto CD-ROM in the 1990s. It may even sell for less than the Pounds sterling 1500 the hard copy cost you, says the publisher. The planned CD-ROM will certainly come with software that can search the entire dictionary in less than a second.
And what will the third edition look like? Hughes says that all the scientific entries need thorough checking. Julia Swannell, a lexicographer, adds that it will be necessary to go over the 19th-century science and add more modern supporting quotations. Although someone will frequently query a particular entry, many scientific words and meanings have not been checked since they were first edited. The potential problems grow worse the earlier in the alphabet a word sits: the dictionary originally appeared in alphabetical order over nearly 50 years, and some words have not been checked for a century. Another problem is that tags to describe the context in which a word is used are not always consistent or clear 鈥 in the first edition, for example, 鈥楶hys鈥, stood for 鈥榩hysiology鈥; the Supplement used 鈥楶hysiol鈥; the OED2 uses both. This is no great worry for a printed work, where the context is normally obvious, but could wreak havoc with an electronic system that was asked, say, to count the number of words in the dictionary to do with physiology.
Hughes also faces a potential explosion of words. No one can predict what science will come up with, but genes are already starting to worry him, lexicographically speaking. 鈥業n general, we have ignored genes as potential dictionary items, since they are names.鈥 Like taxonomic names, place names and personal names, genes have been excluded. Until now, the typeface, italic or initial capital letter, has always given a name away. Recently, however, Hughes has begun to notice the italic being dropped from the names of genes. His highlighted copy of Nature for 15 December 1988, for example, shows 鈥榡un鈥 and 鈥榝os鈥 in Roman type. The prospect is worrying: who knows how many genes there are?
How software managed to capture the English language
FROM a distance, the idea of computerising the OED seems only common sense. Once held as digital data, the dictionary could be produced in any form: a book, a compact disc, an online database or any future development in electronics. Equally important, computerisation would open up the storehouse of knowledge that the dictionary represents.
Computerisation also changes a multi-volume work such as the OED into one volume. Even if it is printed in 20 volumes, it can be searched electronically as one. It was, says James Howes, chief computer systems designer at the Oxford University Press, something that 鈥榦ught to be done鈥. But, he adds, the question was 鈥楥ould it really be done?鈥 In the event, it took the equivalent of 14 people working for one year to design and build the system. Like others involved in the project, Howes seems to take the greatest satisfaction not from having completed the task, but from having done it on time and on budget.
A reference work seems an ideal candidate for computerisation. But Howes sees it differently: 鈥楥ommercial computers and the programs written for them are designed around accounting, or inventories, where you define fixed-length fields, or if variable, not very.鈥 The dictionary is built around text entries which vary from 12 bytes to 500K. This, says Howes, is why the University of Waterloo in Canada was keen to take part, because the project allowed researchers in computer science to become involved in a 鈥榬eal world problem鈥 鈥 highly complex text in vast quantities.
The project was simple in outline: key the text into a computer; develop the software to analyse the structure of the OED and the Supplement; use the software to integrate the two works into one; and develop a system so that lexicographers could edit and add to the text.
When James Murray began the dictionary in 1879, he developed his own structure for entries: headwords, variant forms, etymology, senses, quotations, cross-references and so on 鈥 altogether, 41 main elements in the text. Nonetheless, the structure was clear enough, despite the complexity, for a program to handle. Working with the University of Waterloo, the team at Oxford developed its own software to analyse an entry, determine which element was which and apply a descriptive tag, such as 鈥渁uth鈥, 鈥渜uot鈥 or 鈥渆tym鈥. As Howes says, there was 鈥榥o solution sitting for us on the shelf鈥.
In software terms, they used finite state automata. A finite state automaton is a tool that assigns a value to an element, on the basis that in such and such a context an element can be one, and only one, of a finite number of states. Their programs went through the entire dictionary, determining each element from its position in the text and its surrounding elements.
鈥楢ctually, it鈥檚 easy to find a rule to describe any text,鈥 said Howes. 鈥楤ut is it useful? Does it correspond to the real world?鈥 In the case of the dictionary, the rules had to correspond to the rules implicit in the text, or explicit only in the form of typography (as bold type or upper case, for example). 鈥榃e came up with a reasonable approximation,鈥 says Howes with delicate understatement.
The test was whether the program could integrate the OED with its Supplement. In the event, the software team had to run the dictionary through several passes, as well as develop two 鈥榙ialects鈥, one for the OED and one for the Supplement. At the end of the process, everything was identified and labelled. Just as importantly, the structure of the dictionary, with up to nine nested layers in each entry, was held in memory.
In a conventional business database, entries are built up from fields. For the dictionary, there is only one field for each text entry: the text is the field. But the system of tagging ensured that software could build an entry鈥檚 structure simply by 鈥榩arsing鈥 it, analysing its grammar.
The next task for the software was to integrate the OED and the Supplement. As many of the entries in the Supplement were only additions to entries in the main dictionary, the integration left a problem with cross-references 鈥 600 000 of them. The team designed a number of programs to check whether the cross-references were still valid, and, if not, renumber them.
Software even dealt with pronunciation. Murray used his own highly consistent system to illustrate how a word sounds. The editors of the new dictionary thought it would be better to use the International Phonetic Alphabet, but the chore of altering hundreds of thousands of phonetic spellings by hand was too daunting.
Again, the team came up with programs to deal with this, automatically converting Murray鈥檚 phonetic spellings to the IPA standard. This was trickier than it might appear, because the value of individual characters often depended on the characters around them, so it was not a simple case of changing one character for another. The program had first to decide which character belonged with which.
The final task for the computer group was to design a system to allow lexicographers to retrieve and edit text. The solution went against the trend of WYSIWYG (what you see is what you get) which now dominates desktop publishing. 鈥榃e were interested in the underlying integrity of the structure,鈥 says Howes.
The problem is that it is very easy to get lost in a text entry of 250K (around 20 times the size of the main article here). The solution was an 鈥榠ntelligent鈥 system that rebuilt the inherent structure of an entry when an editor called it up, and then relied on the structure to move around within the entry. The editors, meanwhile, could suppress various levels of text in order to see more clearly the structure of the whole entry.
Julia Swannell, a lexicographer at the Oxford University Press, senses that a slight lull has set in since the hectic development that culminated in the launch earlier this year. But, she says, 鈥榃e now have the chance to design an electronic text for a really long-term future.鈥