Skip to main content

Too little?

The announcement that undergraduate physics teaching is to be phased out at the University of East Anglia (UEA) has sent a shiver through the UK physics community. Students admitted last year will finish their course, hopefully graduating in 2000, but there will be no fresh intake into physics this autumn, and staff who retire will not be replaced. However, the remaining staff will form a graduate school in physics to carry out research in semiconductors and liquids.

UEA is the fourth physics department to find itself in serious difficulty this academic year. De Montfort University in Leicester and Coventry University have already announced that they are closing their departments of applied physics, while physicists at Birkbeck College in London are fighting a decision to close the physics department. Other departments are rumoured to be facing similar predicaments.

Birkbeck is something of a special case in that, uniquely among British universities, it offers only part-time undergraduate degrees, although it admits full-time postgraduates and has a long tradition in research. It could be argued that De Montfort and Coventry are also special cases – they are former polytechnics and did not enter their physics departments for the last research assessment exercise (RAE). But UEA is a traditional university, albeit one established as recently as 1963. Moreover, the eight research active staff at UEA received a ‘4’ in the last RAE – that is, national excellence in all areas and international excellence in some.

The four departments have two things in common: they are small departments and they are in deficit to their universities. The minimum feasible size for a physics department has long been a subject of debate. Frequently this debate has centred on the number of academic staff, but in these four cases it is the low number of student admissions – and the resulting loss of £3000 or so per year that each full-time student brings to the university – that has been the problem. This shortage of students is one of the main reasons why the departments are running a deficit. However, this problem is not restricted to smaller departments. An informal survey of 27 departments two years ago found that 20 were in the red, some to the tune of £20000 per staff member per year. For even a small department of just 10 staff, this could amount to an annual deficit of £200000, something that most bursars are unlikely to look kindly on, especially if it shows no signs of improving.

Many reasons have been advanced for these deficits – mostly to do with physics departments receiving a low “unit of resource” from the higher education funding councils and an inadequate overhead on grants from the research councils. Physics is feeling the pinch as more and more universities insist on “balancing their books” on a department by department basis. But the most worrying long-term trend is the low numbers of students choosing to study physics at universities. Numbers have been stable around 2900 for the past three years, but with total student numbers increasing, and the funding for each student falling, this status quo has worked against physics. The decision at UEA was forced by admissions falling from 25-30 in the past to around 18 last year and the year before.

If we look at schools, the number of students in England and Wales sitting A-levels, a pre-requisite for admission to most universities, in physics has dropped from 45000 in 1988 to 33000 in 1996, a fall of 28%. A lack of good physics teachers, a perception that physics is difficult, doubts about careers in physics, and a wider range of choices at A-level have all worked against the subject.

What can be done? The Institute of Physics is launching a £1m campaign to bring the 16-19 syllabus up to date, efforts to convey the excitement and relevance of science to the public continue, and the Teacher Training Agency is doing its best to recruit more science teachers. Applications were up by a third last year – one can only hope that all these initiatives are not too late.

Access all areas

This is the 100th issue of Physics World. The occasion is marked by a small section that starts with an article by Philip Campbell, who edited the first 85 issues of Physics World before leaving to become editor of Nature, and ends with my own efforts to predict the future. In that item I make a promise that articles in future issues of the magazine will be easier to understand. Quite by coincidence, the 12 December issue of Nature contains an editorial entitled ‘In pursuit of comprehension’ which outlines plans to inject ‘increased effort into the readability of its papers’. It is almost as if, after ten years of fretting about the need for better public understanding of science (and engineering and technology), there has been belated and sudden recognition of the need for improved understanding among scientists themselves. For Physics World, this involves reporting developments throughout physics to all sorts of physicists in as accessible, authoritative and timely a manner as possible. The need for accessibility is most acute in the “features” and “physics in action” sections of the magazine. To ensure authority, these articles are always written by acknowledged experts in the field who can put the work in context by highlighting what is new and why it is important. Physics World staff work closely with the authors during the editing stage to ensure that the article is accessible. (Timeliness can be difficult in a monthly magazine in which many of the articles are written by busy researchers but, I feel, we all do our best).

Spelling out acronyms and explaining jargon can aid accessibility, but there is one unavoidable, and difficult to answer, question: what can we assume our readers know or remember? It would be impossible to compile a list that answers this question but we assume, for example, that there is no need to explain what a crystal lattice is, that it does no harm to add that phonons are vibrations of the lattice, and that is it absolutely essential to explain what a Brillouin zone is every time one is mentioned. Ionization is well known but variants like auto-ionization and above-threshold ionization are always explained. And we try not to patronize you by calling colliders “atom smashers” or saying that something is “thinner than a human hair”.

Despite these efforts, this issue contains several articles that are not for the faint hearted. The “physics in action” section, for example, contains articles that report on the observation of glass-like behaviour in proteins, and the first experiment to measure the decoherence of quantum wavefunctions induced by the environment. However, if the protein article, for example, gives the reader some inkling of the amount of difficult physics currently being studied in proteins, and if the definitions of proximal and distal haemopockets are enough to get you through the article, then that is a start.

The Physics World editorial staff have felt the need for improved accessibility for some time, although the response to our recent reader survey suggests that fine tuning rather than wholesale change is needed. But we do not underestimate the improvements that need to be made if we are to meet the challenge of making the latest breakthroughs in biophysics, quantum theory and other specialities accessible to readers outside these fields.

In the 1960s Bob Dylan sang “don’t criticize what you can’t understand”. The fact that so many of the readers who replied to the survey did not feel the need to criticize the magazine could mean that they are Dylan fans who do not understand the magazine, or that they understand it perfectly. Of course this is an exaggeration, but it is comforting to us to know that the truth is closer to the latter than the former, and comforting to readers, I hope, that the magazine will continue to move in this direction.

Electronic publishing and visions of hypertext

Will editors of journals and magazines such as this be out looking for new jobs in a few years’ time? Will a world overrun with forests use paper only for packing the confectionery eaten by hungry hackers? Should you save this issue of Physics World as a possible collector’s item?

Experience with computer networks, and in particular with the “World-Wide Web” (“W3”) global information initiative (see box “The World-Wide Web”), suggests that the whole mechanism of academic research will change with new technology. But when we try (dangerously) to envisage the shape of things to come, it seems that some old institutions may resurface, albeit in a new form.

The change from paper to electronic form is, perhaps most significantly, a change of timing. It will take the same amount of time to read a page of text, but to follow up a reference will take a few seconds rather than a few days. It will take the same amount of time to compose an article, but to search a library catalogue will take a few seconds rather than a few hours (including the trip to the library). The change in timing will affect the whole way we do work.

Picture a scenario in which any note I write on my computer I can “publish” just by giving it a name. In that note I can make references to any other article anywhere in the world in such a way that when reading my note you can click with your mouse and bring the referenced article up on your screen. Suppose, moreover, that everyone has this capability.

The World-Wide Web

The World-Wide Web (more easily pronounced as “W3”) is an initiative to allow any information on the networks to be easily shared by non-specialists. W3 was conceived at CERN as an essential infrastructure for the particle physics community, but its application has spread into many disciplines.

Technically, W3 uses hypertext (see box “What is ‘hypertext’ anyway?”), text retrieval and wide-area networking techniques. The user runs a browser programme on his/her local machine: this programme hides all the technical details, network connections, data format and protocols. The web is completely distributed: anyone can “publish” data and make links to other documents.

To the user, W3 gives a simple consistent point-and-click or command line interface to a vast wealth of information. Servers exist at several particle physics institutes, and we encourage new institutes to join the web.

All information about W3 is of course available by browsing the web, but questions may be mailed to: www-bug@info.cern.ch (or from JANET, www-bug@ch.cern.info).Those on Internet who wish to try it out in its simplest form can telnet info.cern.ch (no username/password). Better, pick up a browser by anonymous FTP from info.cern.ch.

These are the assumptions of “global hypertext” (see box “What is ‘hypertext’ anyway?”) and it is generally supposed that this will lead to a tangled web of interconnected jottings representing the sum of human knowledge. By selectively following links passed to me by friends, I can rapidly find anything I want to know. Modelling the real world with all its random associations, the “web” allows me to replace a day’s worth of library visits, discussions over coffee and rummaging in filing cabinets with a dozen or so clicks with the mouse.

Before readers jump in and take this to pieces, let me assure the cynical that to a certain extent, this exists, and where it exists, it works. The W3 initiative at CERN and various other institutes has, along with a few similar projects, put together the infrastructure of network protocols and common software. It seems to be taking off, to judge from the readership of our own “server” doubling every other month and new servers cropping up increasingly frequently. Even without global authorship, global readership of data provided by the few has been spectacularly successful. Almost all the data on the web are a window onto some other source, so it is not hand-crafted hypertext, but it’s in demand nevertheless.

Tim Berners-Lee with Nicola Pellow

In this happy anarchy, two problems arise. One is that of collective schizophrenia. The bulk of human understanding may well develop two independent pockets of knowledge about the same thing. This can happen on a small scale, when one writes a document with the sinking feeling that one has written it before but can’t find it. It can happen on a global scale when researchers on different continents investigate the same phenomenon, unaware of each other.

To solve this, some global co-ordination is clearly required. However, centralised co-ordination is out of the question for an estimated 1014 documents. A number of people have started to make lists of resources on the network and have generally been swamped by its growth. The most spectacular success is the “archie” project which keeps a mammoth index of the names of almost all the files available in the internet archives worldwide. Even Peter Deutch, its Canadian instigator, admits that network information is likely to grow faster than his disks, and that his indexes will have to become specialised.

My own attempt to edit a hypertext encyclopaedia, in which pointers to network information sources are classified by subject, leaves me overwhelmed even now. As I looked around for people to help, I realised that I was looking for specialists in particular fields to look after them – like specialised librarians.

What is 'hypertext' anyway?

The term “hypertext” was coined by Ted Nelson, something of a guru in this field, in the 1950s, but only recently has wide-area hypertext become a reality, by allowing one to follow references (“links”) at will.

Typically, one clicks with a mouse on a highlighted phrase, and another related document is displayed. Hypertext is both a new medium for writing, and also a convenient way of representing existing multiply connected information.

The bringing together of the provider of information and the enquirer, the “resource discovery” problem, is up for grabs in the networking community. Solutions, however, always centre on some idea of “subject”. The keyword list, or vocabulary profile, of a document is used to route it to some specialised index which will note it, and direct enquirers to it. Whether you take the Dewey decimal system or the English language as a basis, there need to be centres of knowledge on particular subjects. So, at CERN, we keep pointers to information at other high-energy physics sites. I’d like to see more of this. I’d like someone to maintain an eminently readable hypertext overview of the field, with links to more detailed discussions of specific areas, and eventually to the work of particular groups and individuals. I am not sure whether I would call the result an encyclopaedia, or a journal, or a library. The job-title “cybrarian” has been suggested. However, I can tell you from experience that it takes an incredible amount of time. The tasks of librarians and reviewers are not going to be usurped by academics in their spare time.

The second problem which the information web faces really started with the laser printer. Before laser printers came around, you could tell something about the reliability of an article by its feel: handwritten scrawl torn off a spiral notebook never carried quite the authority of glossy typography. Nowadays, it all looks the same, elegant Optima 10 point.

The same will be true of networked information. It is true that individuals slip into new conventions for conveying formality or lack thereof. A lower case i for the first person gives electronic mail a “hastily scribbled” impression, not to mention the conventional faces on their side (:-) of Internet news. As well as conventions, new ethics for electronic publishing are developing. However, one needs the equivalent of a refereed journal to convey authority. In the hypertext world, the actual physical distribution of data is not the issue here: it is the organisation. The feel of the paper of a document will be replaced by its registered name. A document registered under my personal authorship will not carry the same weight as one registered in the International Standards Organisation’s catalogue of standards.

I see the need for two organs: the newsletter and the encyclopaedia. An encyclopaedia will be an overall attempt by the knowledgeable, the learned societies or anyone else, to represent the state-of-the-art in their field. An encyclopaedia will be a living document, as up to date as it can be, instantly accessible at any time. It will contain carefully authored explanations and summaries of the subject, as well as computer-generated indexes of literature. A reference to a paper from the encyclopaedia conveys authority and acceptance by academic society. A measure of a paper’s standing may be conveyed by the number of links it is away from an encyclopaedia.

I see the need for two organs: the newsletter and the encyclopaedia.

Tim Berners-Lee, writing in 1992

The newsletter is a commentary on the changes in the field. A personalised newsletter can be generated automatically by looking for changes in the encyclopaedia and linked works. A user may effectively ask his computer each day, “Tell me anything new which has been linked to one of my favourite subjects”. It may be possible to generate a newsletter largely automatically, but a human being does so much better. This is especially true of the job of summarising in a review or contents page.

screenshot showing the World-Wide Web browser created by Tim Berners-Lee

I say “an encyclopaedia” rather than “the encyclopaedia”. Another fundamental change will be the low start-up cost of publishing. Anyone can start a new encyclopaedia, and, if enough people refer to it, it will be widely read, and quoted by society’s established authorities. This allows for many encyclopaedias, even many parallel societies. Conventional science will have no hold on pockets of alternative ideas and, so long as innocent individuals are not misled, I see great worth in this freedom. Here we hope to see a market economy in information. The quality of an article is judged by its own contents, and by the quality of the articles to which it refers. There is therefore an incentive to refer to good articles, so the better articles will be most referred to, and most read. Authoritative sources will take care only to refer to reliable work, but deserving small journals can start and rapidly gain prestige. The many medium-size discussion groups of the Internet news system provide a vehicle for bringing new sources quickly to light and establishing acceptance.

These arguments convince me that publishing houses, far from being unnecessary, will be in for very exciting times. Their jobs and those of librarians seem to have merged into one as classifiers and reviewers of the world’s knowledge. The practical issue arises of how to pay these good people. There is something distasteful about charging by the byte. The feeling of freedom to browse is marred by a price tag on an icon, a taxi-meter ticking away on the corner of the screen. Also, it is difficult to find a general method for previewing a document so that one can see what one will get for one’s money. How can one ascertain which documents were really read, or even then were actually useful to the reader? The clear advantage of this technique, however, is that the information “market” becomes more real and more direct when real money is tied in at a low level.

Payment by subscription also has its appeal. Just as one subscribes to a journal, or gains the right to use a library, so one would subscribe to an information service – paying not for the information read but for that which is available just in case. However, the power of global hypertext to represent knowledge lies in the unconstrained way links can cross boundaries between organisations, subjects, and continents. Following a link should ideally take under a quarter of a second (so as not to disturb the train of thought). It should not be accompanied by questions about account numbers and credit ratings.

Following a link should ideally take under a quarter of a second. It should not be accompanied by questions about account numbers and credit ratings.

Tim Berners-Lee, writing in 1992

As a third possibility, the charging and the paying may be done between organisations, over negotiating tables, behind the back of the poor researcher. Let a consortium of physics institutes commission an electronic journal, give it a budget and review it from time to time, cross-licensing between societies so that, for example, members of a national physics society will be granted access to chemistry journals. Perhaps we can imagine an association of publishers which attempts to redistribute the money in a fair fashion.

A mixture of such schemes may exist, and the market may decide which one works best. The market will be fierce, and enthusiastic amateurs will always be willing to compete where they feel a professional service falls below a certain standard.

The existing web gives a good feel for what is possible, using free information. Within high-energy physics, the web contains mainly user manuals, online help, phone books, discussion lists, announcements, news, the minutes of meetings and preprint lists. In other subjects, data range from catalogues of DNA sequences and chemical formulae, through poetry, prose and religious books to the weather forecast. Thanks to the spread of the Internet, this is available in most academic institutes. Already it shows us a more efficient way to pool our knowledge, while keeping up standards of freedom of information which academics, and the Internet, have always promoted.

Rudolf Peierls: In defence of ‘measurement’

Rudolf Peierls lecturing at a blackboard

In a stimulating article (“Against ‘measurement’” Physics World August 1990) the late John Bell professed dissatisfaction with the foundations of quantum mechanics as usually presented, particularly in connection with the so-called “collapse of the wavefunction” as a result of a measurement. He agreed that for all practical purposes the use of quantum mechanics by qualified practitioners leads to well defined answers which, where they can be checked, agree with experiment.

However, he regarded it as necessary to have a clearly formulated presentation of the physical significance of the theory without relying on ill-defined concepts. I agree with him that this is desirable, and, like him, I do not know of any textbook which explains these matters to my satisfaction. I agree in particular that the books he quoted do not give satisfactory answers (I assume that they are fairly quoted; I have not re-read them).

But I do not agree with John Bell that these problems are very difficult. I think it is easy to give an acceptable account, and in this article I shall try to do so. I shall not aim at a rigorous axiomatic, but only at the level of the logic of the working physicist.

The most fundamental statement of quantum mechanics is that the wavefunction, or more generally the density matrix, represents our knowledge of the system we are trying to describe

Rudolf Peierls

In my view the most fundamental statement of quantum mechanics is that the wavefunction, or more generally the density matrix, represents our knowledge of the system we are trying to describe. I shall return later to the question “whose knowledge?”. It is well known that we have to use a wavefunction if we have a “pure state” i.e. if our knowledge of the system is complete, in the sense that any further knowledge is barred by the uncertainty principle. Failing such complete knowledge we must use a density matrix, which therefore contains both quantum and classical ignorance. The wavefunction is a special case of a density matrix, and I shall here talk about “density matrix” when I mean “wavefunction or density matrix”.

More precisely, while the time variation of the density matrix is given by Schrödinger’s equation, the initial values represent knowledge usually obtained from observations. (There are not always measurements; for example, if an atom has been for a reasonable time in free space we know it must be in its ground state.)

In quantum mechanics we have to be specific about what we know, because our possible knowledge is confined by the uncertainty principle

Rudolf Peierls

Our knowledge is not fixed, but may increase or decrease. It increases if further observations are made; it decreases if the system is disturbed by external factors which we cannot control. There is nothing new in this. In classical physics our knowledge may increase and decrease in the same way. The only difference is that in quantum mechanics we have to be specific about what we know, because our possible knowledge is confined by the uncertainty principle. In classical physics there is no reason in principle why we cannot know everything about the system and we usually argue as if we did. But in a practical situation our knowledge may increase or decrease as indicated.

In quantum mechanics any increase in our knowledge is usually accompanied by a decrease in some other respect, because of the uncertainty principle. This applies particularly when we are concerned with a “pure state”. Then we can gain no new information (other than confirming what we know already) without losing some of the existing information.

Once this significance of the density matrix is understood, it is clear that upon a change in our knowledge the density matrix must change. This is not a physical process, and we certainly cannot expect it to follow from the Schrödinger equation. It is just the fact that our knowledge has changed, and thus must be represented by a new density matrix.

When I refer to “observation”, this term has its common-sense meaning. The observation usually (but not necessarily) involves an apparatus which interacts with the system in question, and which produces a signal (visible, audible, or other) which we can recognise, and which is correlated with the variables of the system. Bell quoted the view of Landau and Lifshitz (and therefore of Bohr) that the apparatus must necessarily obey classical physics. In my view this is not correct. It is of course true that our senses are macroscopic, and that the instruments we find convenient are also macroscopic and in practice classical. But this is a practical point, not one of principle. The sensitivity of the human eye is almost sufficient to detect a single photon. If some experimentalist has sufficient vision to see one photon, the observation of that photon might perfectly well serve as a measurement.

The apparatus usually consists of a chain of correlated events. I have elsewhere (Peierls 1979, 1985) discussed as an example the observation of a spin component of a spin-1/2 atom by a Stern–Gerlach magnet. The first step, the passage through the inhomogeneous magnetic field, sets up a correlation between the spin component and the position of the atom. It is not yet a measurement; we have not yet gathered any information. This requires determining the position of the particle, i.e. in which part of the split beam it travels. To find this out, we may use a counter, but again this conveys no information – and nothing collapses! – until we find out whether the counter has been activated. We can obviously pursue this chain: the counter will be part of an electrical circuit, the circuit will operate a digital recorder, we may read this recorder by means of the light it reflects into our eye, etc. Each step is correlated with the preceding ones and therefore with the spin component of the particle. Each step keeps both options open until we “see” the result, and then we revise our density matrix.

Because of the uncertainty principle we cannot acquire knowledge of, say, the z component of the spin without losing what information we had previously about, say, the x component. Is this happening in the first step, the passage through the magnet? At first sight this looks likely, because information about sx is contained in the phase relation between the components of the wavefunction belonging to sz = +1/2 and sx = –1/2. Since the beams corresponding to the two sz values are now split, they do not overlap and do not interfere, so their phase relationship is not observable. However, the information is not irretrievably lost. By arranging a further magnet we could recombine the two beams and observe their phase relation (thereby foregoing the possibility of observing sz). We do finally lose the “forbidden” information when we “see” the atom in one of the beams. We then have to replace our density matrix by one containing only the one sz value, so there is no interference.

As long as we do not “see” the atom in the beam, the reconstruction of the seemingly lost information is troublesome, but easy to visualise. At the next stage, i.e. after the counter, it becomes much more involved. Since the density matrix now contains the variables of the counter, interference requires not only that the two atomic beams be made to overlap, but in addition that there be an overlap between the density matrices for the activated and unactivated states of the counter. The observation of the phase relation therefore requires an operator capable of deactivating the counter coherently. This is possible in principle, but in practice prohibitively difficult. As we go further down the chain of connections involved in our “measurement”, this difficulty gets worse.

This is the origin of the belief that the apparatus makes the off-diagonal matrix elements of the density matrix disappear. In most cases that is true “for all practical purposes”, but not in principle. The off-diagonal matrix elements disappear only when we know the result of the measurement.

The “system” to which we apply our description can be as large as we like, including the whole world if we want. However, if we make the system too large, the amount of information we can obtain is relatively small, so that the density matrix is made up mostly of parts proportional to the unit matrix (which denotes complete ignorance) and it becomes hard to do any useful physics. In any case the “system” cannot include the mind of the observer and his knowledge, because present physics is not able to describe mind and knowledge (and it is not obvious that this is a proper subject for physics).

The objection is sometimes made: “How can one apply quantum mechanics to the early Universe, when there were no observers around?” The answer is that the observer does not have to be contemporaneous with the event. We can, from present evidence, draw conclusions about the early Universe, the classical example being the cosmic microwave background. In this sense we are observers. If there is a part of the Universe, or a period in its history, which is not capable of influencing present-day events directly or indirectly, then indeed there would be no sense in applying quantum mechanics to it.

That leaves the question: whose knowledge should be represented in the density matrix? In general there will be many who may have some information about the state of a physical system. Each of them has to use his or her density matrix. These may differ, as the nature and amount of knowledge may differ. People may have observed the system by different methods, with more or less accuracy; they may have seen part of the results of another physicist. However, there are limitations to the extent to which their knowledge may differ. This is imposed by the uncertainty principle. For example if one observer has knowledge of sz of our Stern–Gerlach atom, another may not know sx, since the measurement of sx would have destroyed the other person’s knowledge of sz, and vice versa. This limitation can be compactly and conveniently expressed by the condition that the density matrices used by the two observers must commute with each other.

I must confess that the scheme, with both hidden variables and probability rules, seems to me exceedingly ugly, but of course one cannot argue about this

Rudolf Peierls

John Bell referred to two alternative interpretations of quantum mechanics, that of de Broglie–Bohm (BB), and that of Ghiradi–Rimini–Weber (GRW). As far as I know the BB scheme reproduces all predictions of quantum mechanics. A decision can therefore be made only on aesthetic grounds. I must confess that the scheme, with both hidden variables and probability rules, seems to me exceedingly ugly, but of course one cannot argue about this. I have not studied the implications of the GRW scheme in detail, but I believe that there must be cases where it makes predictions differing from those of quantum mechanics, which would be observable in principle.

Further reading

J S Bell 1990 “Against ‘measurement’” Physics World August 33–40

R Peierls 1979 Surprises in Theoretical Physics Princeton section 1.6 (Some of the points made in this article will also be discussed in a forthcoming volume, More Surprises in Theoretical Physics, Princeton)

R Peierls 1985 “Observations in Quantum Mechanics and the ‘Collapse of the Wave Function’” in Symposium on the Foundations of Modern Physics World Scientific

Isaac Asimov: ‘There is nothing more fascinating than odd corners of mathematics’

To some minds, including that of the reviewer, there is nothing more fascinating than the snapshots of odd corners of the world of mathematics, corners that are made clear to the non-mathematician. For that you need a clever guide, and there are surely few who are cleverer than Malcolm Lines.

In this book (not his only one) he deals with such old familiar matters as the Fibonacci numbers (sequences of integers, where each is the sum of the two preceding it) and Euclid’s fifth postulate, but manages to bring out points of interest in both. He shows the relationship of the Fibonacci numbers to pine cones and the relationship of Euclid’s fifth postulate to the shape of the Universe, and does so with charm and finesse.

He also discusses those parts of mathematics that have made the headlines in recent years. There are the fractals that deal with curves that possess non-integral dimensions, not quite two-dimensional but certainly more than one-dimensional, for instance. And there is the new mathematics of chaos in which the impossibility of setting the original conditions exactly means that consequences become unpredictable, even where apparently quite simple problems are involved.

But there is no point in repeating the entire book. Let me instead make two points that have risen forcibly to mind as a result of reading it.

First, is the matter of elegance in mathematics. Suppose we are faced with a safe that can be opened by a certain combination of letters and numbers. It may be possible by pure reason to work out a combination that is likely to have been chosen by the person who owns the safe. (I have written mystery stories of this sort.) It is very satisfactory to say, ‘Aha, try this combination’. It is tried and the safe falls open.

On the other hand, it is also possible to make use of a gob of explosive, strategically placed, set it off, and, boom, the safe is open. Now the explosive has solved the problem, yet how unsatisfying it is. The method of explosion may work, but it is not elegant.

Well, in mathematics, there is the four-colour problem. Can any map be coloured in four different colours so that no two adjacent regions have the same colour? Mathematicians were reasonably certain this was so. No map had been devised that required more than four colours. On the other hand, to prove that four colours is always sufficient has eluded some very clever mathematicians.

Lines gives the history of the four-colour problem, including some ingenious proofs that invariably turned out to have subtle flaws that invalidated them. Eventually, it boiled down to dealing with some 1500 different configurations, one by one, to show that four colours were sufficient in each case. Finally, in July 1976, the problem was solved – by the use of two or three computers and 1000 hours of time.

The solution is certainly valid, but it is long and incredibly involved and you have to depend on the computers. It’s not elegant and mathematicians are unhappy and disappointed. They want something clever and more insightful – and who knows, they may yet find it.

Second, is the question of utility. Even the most arcane mathematical byways turn out to be useful. The four-colour theorem should be ideal for the ‘Who would care but a mathematician?’ category, but it turns out to have applications to airline schedules and telephone connections. Prime numbers, which have always been thought of as the purest of pure mathematics turn out to be the key to unbreakable cryptography and are therefore essential to national defence.

There’s a certain ‘dirtiness’ about such utility, a certain diminution of an ethereal ideal. What a relief, then, to turn to ‘hailstone numbers’, lovingly described by Lines. You simply pick a number, if it is odd, you multiply it by three and add one; if it is even, you divide it by two. You repeat this with every new number you get and the result is that the numbers rise and fall like hailstones forming in the air, but, eventually, die down to 1.

Different numbers take different lengths of time to die down and, in the process manage to reach different heights. Small numbers don’t usually last long, but 27 goes through 111 steps before dying and, in the process, reaches a height of 7288 as the 67th number and 9232 as the 77th number. No number does better till 255 is reached. It doesn’t last as long before dying but reaches a peak value of 13,120. The number 77,671 reaches a peak of 1570,824,736 before eventually shrinking.

Must all numbers subjected to this rule eventually die? Is there no number that can dance about in the upper air forever? Mathematicians think they must all die, for certainly all numbers less than one trillion (1 × 1012) do – but there is no general proof of the matter.

What interests me, however, is that surely this is something that can have no conceivable use. It is pure game, pure fun!

  • 1990 Adam Hilger 163pp £7.50pb

Against ‘measurement’: John Bell on our continuing struggles with quantum mechanics

Surely, after 62 years, we should have an exact formulation of some serious part of quantum mechanics? By ‘exact’ I do not of course mean ‘exactly true’. I mean only that the theory should be fully formulated in mathematical terms, with nothing left to the discretion of the theoretical physicist … until workable approximations are needed in applications. By ‘serious’ I mean that some substantial fragment of physics should be covered. Nonrelativistic ‘particle’ quantum mechanics, perhaps with the inclusion of the electromagnetic field and a cut-off interaction, is serious enough. For it covers ‘a large part of physics and the whole of chemistry’ (P A M Dirac 1929 Proc. R. Soc. A 123 714). I mean too, by ‘serious’, that ‘apparatus’ should not be separated off from the rest of the world into black boxes, as if it were not made of atoms and not ruled by quantum mechanics.

The question, ‘… should we not have an exact formulation … ? ‘ , is often answered by one or both of two others. I will try to reply to them: Why bother? Why not look it up in a good book?

Why bother?

Perhaps the most distinguished of ‘why bother?’ers has been Dirac (1963 Sci. American 208 May 45). He divided the difficulties of quantum mechanics into two classes, those of the first class and those of the second. The second-class difficulties were essentially the infinities of relativistic quantum field theory. Dirac was very disturbed by these, and was not impressed by the ‘renormalisation’ procedures by which they are circumvented. Dirac tried hard to eliminate these second-class difficulties, and urged others to do likewise. The first-class difficulties concerned the role of the ‘observer’, ‘measurement’, and so on. Dirac thought that these problems were not ripe for solution, and should be left for later. He expected developments in the theory which would make these problems look quite different. It would be a waste of effort to worry overmuch about them now, especially since we get along very well in practice without solving them.

Dirac gives at least this much comfort to those who are troubled by these questions: he sees that they exist and are difficult. Many other distinguished physicists do not. It seems to me that it is among the most sure-footed of quantum physicists, those who have it in their bones, that one finds the greatest impatience with the idea that the ‘foundations of quantum mechanics’ might need some attention. Knowing what is right by instinct, they can become a little impatient with nitpicking distinctions between theorems and assumptions. When they do admit some ambiguity in the usual formulations, they are likely to insist that ordinary quantum mechanics is just fine ‘for all practical purposes’. I agree with them about that: ORDINARY QUANTUM MECHANICS (as far as I know) IS JUST FINE FOR ALL PRACTICAL PURPOSES.

Paul Dirac next to a blackboard

Even when I begin by insisting on this myself, and in capital letters, it is likely to be insisted on repeatedly in the course of the discussion. So it is convenient to have an abbreviation for the last phrase: FOR ALL PRACTICAL PURPOSES = FAPP.

I can imagine a practical geometer, say an architect, being impatient with Euclid’s fifth postulate, or Playfair’s axiom: of course in a plane, through a given point, you can draw only one straight line parallel to a given straight line, at least FAPP. The reasoning of such a natural geometer might not aim at pedantic precision, and new assertions, known in the bones to be right, even if neither among the originally stated assumptions nor derived from them as theorems, might come in at any stage. Perhaps these particular lines in the argument should, in a systematic presentation, be distinguished by this label – FAPP – and the conclusions likewise: QED FAPP.

I expect that mathematicians have classified such fuzzy logics. Certainly they have been much used by physicists. But is there not something to be said for the approach of Euclid? Even now that we know that Euclidean geometry is (in some sense) not quite true? Is it not good to know what follows from what, even if it is not really necessary FAPP? Suppose for example that quantum mechanics were found to resist precise formulation. Suppose that when formulation beyond FAPP is attempted, we find an unmovable finger obstinately pointing outside the subject, to the mind of the observer, to the Hindu scriptures, to God, or even only Gravitation? Would not that be very, very interesting?

But I must say at once that it is not mathematical precision, but physical, with which I will be concerned here. I am not squeamish about delta functions. From the present point of view, the approach of von Neumann’s book is not preferable to that of Dirac’s.

Why not look it up in a good book?

But which good book? In fact it is seldom that a ‘no problem’ person is, on reflection, willing to endorse a treatment already in the literature. Usually the good unproblematic formulation is still in the head of the person in question, who has been too busy with practical things to put it on paper. I think that this reserve, as regards the formulations already in the good books, is well founded. For the good books known to me are not much concerned with physical precision. This is clear already from their vocabulary.

Here are some words which, however legitimate and necessary in application, have no place in a formulation with any pretension to physical precision: system, apparatus, environment, microscopic, macroscopic, reversible, irreversible, observable, information, measurement.

The concepts ‘system’, ‘apparatus’, ‘environment’, immediately imply an artificial division of the world, and an intention to neglect, or take only schematic account of, the interaction across the split. The notions of ‘microscopic’ and ‘macroscopic’ defy precise definition. So also do the notions of ‘reversible’ and ‘irreversible’. Einstein said that it is theory which decides what is ‘observable’. I think he was right – ‘observation’ is a complicated and theory-laden business. Then that notion should not appear in the formulation of fundamental theory. Information? Whose information? Information about what?

On this list of bad words from good books, the worst of all is ‘measurement’. It must have a section to itself.

Against ‘measurement’

When I say that the word ‘measurement’ is even worse than the others, I do not have in mind the use of the word in phrases like ‘measure the mass and width of the Z boson’. I do have in mind its use in the fundamental interpretive rules of quantum mechanics. For example, here they are as given by Dirac (Quantum Mechanics Oxford University Press 1930):

‘. . . any result of a measurement of a real dynamical variable is one of its eigenvalues . . . ”

‘. . . if the measurement of the observable . . . is made a large number of times the average of all the results obtained will be . . .’

‘. . . a measurement always causes the system to jump into an eigenstate of the dynamical variable that is being measured . . .’

It would seem that the theory is exclusively concerned about ‘results of measurement’, and has nothing to say about anything else. What exactly qualifies some physical systems to play the role of ‘measurer’? Was the wavefunction of the world waiting to jump for thousands of millions of years until a single-celled living creature appeared? Or did it have to wait a little longer, for some better qualified system . . . with a PhD? If the theory is to apply to anything but highly idealised laboratory operations, are we not obliged to admit that more or less ‘measurement-like’ processes are going on more or less all the time, more or less everywhere? Do we not have jumping then all the time?

The first charge against ‘measurement’, in the fundamental axioms of quantum mechanics, is that it anchors there the shifty split of the world into ‘system’ and ‘apparatus’. A second charge is that the word comes loaded with meaning from everyday life, meaning which is entirely inappropriate in the quantum context. When it is said that something is ‘measured’ it is difficult not to think of the result as referring to some pre-existing property of the object in question. This is to disregard Bohr’s insistence that in quantum phenomena the apparatus as well as the system is essentially involved. If it were not so, how could we understand, for example, that ‘measurement’ of a component of ‘angular momentum’ – in an arbitrarily chosen direction – yields one of a discrete set of values? When one forgets the role of the apparatus, as the word ‘measurement’ makes all too likely, one despairs of ordinary logic – hence ‘quantum logic’. When one remembers the role of the apparatus, ordinary logic is just fine.

In other contexts, physicists have been able to take words from everyday language and use them as technical terms with no great harm done. Take for example, the ‘strangeness’, ‘charm’, and ‘beauty’ of elementary particle physics. No one is taken in by this ‘baby talk’, as Bruno Touschek called it. Would that it were so with ‘measurement’. But in fact the word has had such a damaging effect on the discussion, that I think it should now be banned altogether in quantum mechanics.

The role of experiment

Even in a low-brow practical account, I think it would be good to replace the word ‘measurement’, in the formulation, by the word ‘experiment’. For the latter word is altogether less misleading. However, the idea that quantum mechanics, our most fundamental physical theory, is exclusively even about the results of experiments would remain disappointing.

In the beginning natural philosophers tried to understand the world around them. Trying to do that they hit upon the great idea of contriving artificially simple situations in which the number of factors involved is reduced to a minimum. Divide and conquer. Experimental science was born. But experiment is a tool. The aim remains: to understand the world. To restrict quantum mechanics to be exclusively about piddling laboratory operations is to betray the great enterprise. A serious formulation will not exclude the big world outside the laboratory.

The quantum mechanics of Landau and Lifshitz

Let us have a look at the good book Quantum Mechanics by L D Landau and E M Lifshitz. I can offer three reasons for this choice:

  1. It is indeed a good book.
  2. It has a very good pedigree. Landau sat at the feet of Bohr. Bohr himself never wrote a systematic account of the theory. Perhaps that of Landau and Lifshitz is the nearest to Bohr that we have.
  3. It is the only book on the subject in which I have read every word.

This last came about because my friend John Sykes enlisted me as technical assistant when he did the English translation. My recommendation of this book has nothing to do with the fact that one per cent of what you pay for it comes to me.

LL emphasise, following Bohr, that quantum mechanics requires for its formulation ‘classical concepts’ – a classical world which intervenes on the quantum system, and in which experimental results occur (brackets after quotes refer to page numbers):

‘. . . It is in principle impossible . . . to formulate the basic concepts of quantum mechanics without using classical mechanics.’ (LL2)

‘. . . The possibility of a quantitative description of the motion of an electron requires the presence also of physical objects which obey classical mechanics to a sufficient degree of accuracy.’ (LL2)

‘. . . the ‘classical object’ is usually called apparatus and its interaction with the electron is spoken of as measurement. However, it must be emphasised that we are here not discussing a process . . . in which the physicist-observer takes part. By measurement, in quantum mechanics, we understand any process of interaction between classical and quantum objects, occurring apart from and independently of any observer. The importance of the concept of measurement in quantum mechanics was elucidated by N Bohr.’ (LL2)

And with Bohr they insist again on the inhumanity of it all:

‘. . . Once again we emphasise that, in speaking of ‘performing a measurement’, we refer to the interaction of an electron with a classical ‘apparatus’, which in no way presupposes the presence of an external observer.’ (LL3)

‘. . . Thus quantum mechanics occupies a very unusual place among physical theories: it contains classical mechanics as a limiting case, yet at the same time it requires this limiting case for its own formulation . . . ” (LL3)

‘. . . consider a system consisting of two parts: a classical apparatus and an electron . . . The states of the apparatus are described by quasiclassical wavefunctions Φn(ξ), where the suffix n corresponds to the ‘reading’ gn of the apparatus, and ξ denotes the set of its coordinates. The classical nature of the apparatus appears in the fact that, at any given instant, we can say with certainty that it is in one of the known states Φn with some definite value of the quantity g; for a quantum system such an assertion would of course be unjustified.’ (LL21)

‘. . . Let Φ0(ξ) be the wavefunction of the initial state of the apparatus . . . and Ψ(q) of the electron . . . the initial wavefunction of the whole system is the product Ψ(q)Φ0(ξ). After the measuring process we obtain a sum of the form

∑nAn(q)Φnξ

where the An(q) are some functions of q.’ (LL22)

‘The classical nature of the apparatus, and the double role of classical mechanics as both the limiting case and the foundation of quantum mechanics, now make their appearance. As has been said above, the classical nature of the apparatus means that, at any instant, the quantity g (the ‘reading of the apparatus’) has some definite value. This enables us to say that the state of the system apparatus + electron after the measurement will in actual fact be described, not by the entire sum, but by only the one term which corresponds to the ‘reading’ gn of the apparatus An(q)Φn(ξ). It follows from this that An(q) is proportional to the wavefunction of the electron after the measurement . . .’ (LL22)

This last is (a generalisation of) the Dirac jump, not an assumption here but a theorem. Note, however, that it has become a theorem only by virtue of another jump being assumed – that of a ‘classical’ apparatus into an eigenstate of its ‘reading’. It will be convenient later to refer to this last, the spontaneous jump of a macroscopic system into a definite macroscopic configuration, as the LL jump. And the forced jump of a quantum system as a result of ‘measurement’ – an external intervention – as the Dirac jump. I am not implying that these men were the inventors of these concepts. They used them in references that I can give.

According to LL (LL24), measurement (I think they mean the LL jump)’. . . brings about a new state . . . Thus the very nature of the process of measurement involves a far-reaching principle of irreversibility . . . causes the two directions of time to be physically non-equivalent, i.e. creates a difference between the future and the past.’

The LL formulation, with vaguely defined wavefunction collapse, when used with good taste and discretion, is adequate FAPP. It remains that the theory is ambiguous in principle, about exactly when and exactly how the collapse occurs, about what is microscopic and what is macroscopic, what quantum and what classical. We are allowed to ask: is such ambiguity dictated by experimental facts? Or could  theoretical physicists do better if they tried harder?

The quantum mechanics of K Gottfried

The second good book that we will look at here is that of Kurt Gottfried (Quantum Mechanics Benjamin 1966). Again I can give three reasons for this choice:

  1. It is indeed a good book. The CERN library had four copies. Two have been stolen – already a good sign. The two that remain are falling apart from much use.
  2. It has a very good pedigree. Kurt Gottfried was inspired by the treatments of Dirac and Pauli. His personal teachers were J D Jackson, J Schwinger, V F Weisskopf and J Goldstone. As consultants he had P Martin, C Schwartz, W Furry and D Yennie.
  3. I have read some of it more than once.

This last came about as follows. I have often had the pleasure of discussing these things with Viki Weisskopf. Always he would end up with ‘you should read Kurt Gottfried’. Always I would say ‘I have read Kurt Gottfried’. But Viki would always say again next time ‘you should read Kurt Gottfried’. So finally I read again some parts of KG, and again, and again, and again.

At the beginning of the book there is a declaration of priorities (KG1): ‘. . . The creation of quantum mechanics in the period 1924–28 restored logical consistency to its rightful place in theoretical physics. Of even greater importance, it provided us with a theory that appears to be in complete accord with our empirical knowledge of all nonrelativistic phenomena . . .’

The first of these two propositions, admittedly the less important, is actually given rather little attention in the book. One can regret this a bit, in the rather narrow context of the particular present enquiry – into the possibility of precision. More generally, KG’s priorities are those of all right-thinking people.

The book itself is above all pedagogical. The student is taken gently by the hand, and soon finds herself or himself doing quantum mechanics, without pain – and almost without thought. The essential division of KG’s world into system and apparatus, quantum and classical, a notion that might disturb the student, is gently implicit rather than brutally explicit. No explicit guidance is then given as to how in practice this shifty division is to be made. The student is simply left to pick up good habits by being exposed to good examples.

KG declares that the task of the theory is (KG 16)’. . . to predict the results of measurements on the system . . .’ The basic structure of KG’s world is then W = S + R where S is the quantum system, and R is the rest of the world – from which measurements on S are made. When our only interpretive axioms are about measurement results (or findings (KG11)) we absolutely need such a base R from which measurements can be made. There can be no question then of identifying the quantum system S with the whole world W. There can be no question – without changing the axioms – of getting rid of the shifty split. Sometimes some authors of ‘quantum measurement’ theories seem to be trying to do just that. It is like a snake trying to swallow itself by the tail. It can be done – up to a point. But it becomes embarrassing for the spectators even before it becomes uncomfortable for the snake.

But there is something which can and must be done – to analyse theoretically not removing the split, which cannot be done with the usual axioms, but shifting it. This is taken up in KG’s chapter 4: ‘The Measurement Process . . . ” Surely ‘apparatus’ can be seen as made of atoms? And it often happens that we do not know, or not well enough, either a priori or by experience, the functioning of some system that we would regard as ‘apparatus’. The theory can help us with this only if we take this ‘apparatus’ A out of the rest of the world R and treat it together with S as part of an enlarged quantum system S‘ : R = A + R‘; S + A = S‘; W = S‘ + R‘. The original axioms about ‘measurement’ (whatever they were exactly) are then applied not at the S/A interface, but at the A/R‘ interface – where for some reason it is regarded as more safe to do so. In real life it would not be possible to find any such point of division which would be exactly safe. For example, strictly speaking it would not be exactly safe to take it between the counters, say, and the computer – slicing neatly through some of the atoms of the wires. But with some idealisation, which might ‘. . . be highly stylised and not do justice to the enormous complexity of an actual laboratory experiment . . . ” (KG165), it might be possible to find more than one not too implausible way of dividing the world up. Clearly it is necessary to check that different choices give consistent results (FAPP). A disclaimer towards the end of KG’s chapter 4 suggests that that, and only that, is the modest aim of that chapter (KG189): ‘. . . we emphasise that our discussion has merely consisted of several demonstrations of internal consistency . . . ” But reading reveals other ambitions.

Neglecting the interaction of A with R‘, the joint system S‘ = S + A is found to end, in virtue of the Schrödinger equation, after the ‘measurement’ on S by A, in a state

Ψ=∑ncnΨn

where the states Ψn are supposed each to have a definite apparatus pointer reading gn. The corresponding density matrix is

ρ=∑n∑mcncn*ΨnΨm*

At this point KG insists very much on the fact that A, and so S‘, is a macroscopic system. For macroscopic systems, he says, (KG186) ‘. . . trAρ^ = trAρ for all observables A known to occur in nature . . . ” where

ρ^=∑n|cn|2ΨnΨn*

i.e. ρ^ is obtained from ρ by dropping interference terms involving pairs of macroscopically different states. Then (KG188) ‘. . . we are free to replace ρ by ρ^ after the measurement, safe in the knowledge that the error will never be found . . . ”

Now, while quite uncomfortable with the concept ‘all known observables’, I am fully convinced of the practical elusiveness, even the absence FAPP, of interference between macroscopically different states (J S Bell and M Nauenberg 1966 ‘The moral aspects of quantum mechanics’ in Preludes in Theoretical Physics North-Holland). So let us go along with KG on this and see where it leads: ‘. . . If we take advantage of the indistinguishability of ρ and ρ^ to say that ρ^ is the state of the system subsequent to measurement, the intuitive interpretation of cm as a probability amplitude emerges without further ado. This is because cm enters ρ^ only via |cm|2 , and the latter quantity appears in ρ^ in precisely the same manner as probabilities do in classical statistical physics . . .’

I am quite puzzled by this. If one were not actually on the lookout for probabilities, I think the obvious interpretation of even ρ^ would be that the system is in a state in which the various Ψs somehow coexist: Ψ1Ψ1* and Ψ2Ψ2* and . . .

This is not at all a probability interpretation, in which the different terms are seen not as coexisting, but as alternatives: Ψ1Ψ1* or Ψ2Ψ2* or . . .

The idea that elimination of coherence, in one way or another, implies the replacement of ‘and’ by ‘or’, is a very common one among solvers of the ‘measurement problem’. It has always puzzled me.

It would be difficult to exaggerate the importance attached by KG to the replacement of ρ by ρ^: ‘. . . To the extent that nonclassical interference terms (such as cmc*m) are present in the mathematical expression for ρ . . . the numbers cm are intuitively uninterpretable, and the theory is an empty mathematical formalism . . .’ (KG 187)

But this suggests that the original theory, ‘an empty mathematical formalism’, is not just being approximated – but discarded and replaced. And yet elsewhere KG seems clear that it is in the business of approximation that he is engaged, approximation of the sort that introduces irreversibility in the passage from classical mechanics to  thermodynamics: ‘. . . In this connection one should note that in approximating ρ by ρ^ one introduces irreversibility, because the time-reversed Schrödinger equation cannot retrieve ρ from ρ^.’ (KG188)

New light is thrown on KG’s ideas by a recent recapitulation, referred to in the following as KGR (K Gottfried ‘Does quantum mechanics describe the collapse of the wavefunction?’ Presented at 62 Years of Uncertainty, Erice, 5–14 August 1989). This is dedicated to the proposition that (KGR1) ‘. . . the laws of quantum mechanics yield the results of measurements . . . ” These laws are taken to be (KGR1): ‘(1) a pure state is described by some vector in Hilbert space from which expectation values of observables are computed in the standard way; and (2) the time evolution is a unitary transformation on that vector’ (KGR1). Not included in the laws is (KGR1) von Neumann’s ‘. . . infamous postulate: the measurement act ‘collapses’ the state into one in which there are no interference terms between different states of the measurement apparatus . . . ” Indeed, (KGR1) ‘the reduction postulate is an ugly scar on what would be a beautiful theory if it could be removed . . .’

Perhaps it is useful to recall here just how the infamous postulate is formulated by von Neumann (J von Neumann 1955 Mathematical Foundations of Quantum Mechanics Princeton University Press). If we look back we find that what vN actually postulates (vN347, 418) is that ‘measurement’ ­– an external intervention by R on S – causes the state

ϕ=∑ncnϕn

to jump, with various probabilities into Φ1 or Φ2 or . . .

From the ‘or’ here, replacing the ‘and’, as a result of external intervention, vN infers that the resulting density matrix, averaged over the several possibilities, has no interference terms between states of the system which correspond to different measurement results (vN347). I would emphasise several points here.

  1. von Neumann presents the disappearance of coherence in the density matrix, not as a postulate, but as a consequence of a postulate. The postulate is made at the wavefunction level, and is just that already made by Dirac for example.
  2. I cannot imagine von Neumann arguing in the opposite direction, that lack of interference in the density matrix implies, without further ado, ‘or’ replacing ‘and’ at the wavefunction level. A special postulate to that effect would be required.
  3. von Neumann is concerned here with what happens to the state of the system that has suffered the measurement – an external intervention. In application to the extended system S‘(= S + A) von Neumann’s collapse would not occur before external intervention from R’. It would be surprising if this consequence of external intervention on S‘ could be inferred from the purely internal Schrödinger equation for S‘. Now KG’s collapse, although justified by reference to ‘all known observables’ at the S’/R‘ interface, occurs after ‘measurement’ by A on S, but before interaction across S’/R‘. Thus the collapse which KG discusses is not that which von Neumann infamously postulates. It is the LL collapse rather than that of von Neumann and Dirac.

The explicit assumption that expectation values are to be calculated in the usual way throws light on the subsequent falling out of the usual probability interpretation ‘without further ado’. For the rules for calculating expectation values, applied to projection operators for example, yield the Born probabilities for eigenvalues. The mystery is then: what has the author actually derived rather than assumed? And why does he insist that probabilities appear only after the butchering of ρ into ρ^, the theory remaining an ’empty mathematical formalism’ so long as ρ is retained? Dirac, von Neumann, and the others, nonchalantly assumed the usual rules for expectation values, and so probabilities, in the context of the unbutchered theory. Reference to the usual rules for expectation values also makes clear what KG’s probabilities are probabilities of. They are probabilities of ‘measurement’ results, of external results of external interventions, from R‘ on S‘ in the application. We must not drift into thinking of them as probabilities of intrinsic properties of S‘ independent of, or before, ‘measurement’. Concepts like that have no place in the orthodox theory.

John von Neumann sat in an armchair

Having tried hard to understand what KG has written, I will finally permit myself some guesses about what he may have in mind. I think that from the beginning KG tacitly assumes the Dirac rules at S’/R’ – including the Dirac-von Neumann jump, required to get the correlations between results of successive (moral) measurements. Then, for ‘all known observables’, he sees that the ‘measurement’ results at S’/R’ are AS IF (FAPP) the LL jump had occurred in S‘. This is important, for it shows how, FAPP, we can get away with attributing definite classical properties to ‘apparatus’ while believing it to be governed by quantum mechanics. But a jump assumption remains. LL derived the Dirac jump from the assumed LL jump. KG derives, FAPP, the LL jump from assumptions at the shifted split R’/S’ which include the Dirac jump there.

It seems to me that there is then some conceptual drift in the argument. The qualification ‘as if (FAPP)’ is dropped, and it is supposed that the LL jump really takes place. The drift is away from the ‘measurement’ ( . . . external intervention . . .) orientation of orthodox quantum mechanics towards the idea that systems, such as S‘ above, have intrinsic properties – independently of and before observation. In particular the readings of experimental apparatus are supposed to be really there before they are read. This would explain KG’s reluctance to interpret the unbutchered density matrix ρ, for the interference terms there could seem to imply the simultaneous existence of different readings. It would explain his need to collapse ρ into ρ^, in contrast with von Neumann and the others, without external intervention across the last split S’/R’. It would explain why he is anxious to obtain this reduction from the internal Schrödinger equation of S‘. (It would not explain the reference to ‘all known observables’ ­– at the S’/R‘ split.) The resulting theory would be one in which some ‘macroscopic’ ‘physical attributes’ have values at all times, with a dynamics that is related somehow to the butchering of ρ into ρ^ – which is seen as somehow not incompatible with the internal Schrödinger equation of the system. Such a theory, assuming intrinsic properties, would not need external intervention, would not need the shifty split. But the retention of the vague word ‘macroscopic’ would reveal limited ambition as regards precision. To avoid the vague ‘microscopic’ ‘macroscopic’ distinction – again a shifty split – I think one would be led to introduce variables which have values even on the smallest scale. If the exactness of the Schrödinger equation is maintained, I see this leading towards the picture of de Broglie and Bohm.

The quantum mechanics of N G van Kampen

Let us look at one more good book, namely Physica A 153 (1988), and more specifically at the contribution: ‘Ten theorems about quantum mechanical measurements’, by N G van Kampen. This paper is distinguished especially by its robust common sense. The author has no patience with ‘. . . such mind-boggling fantasies as the many world interpretation . . . ” (vK98). He dismisses out of hand the notion of von Neumann, Pauli, Wigner – that ‘measurement’ might be complete only in the mind of the observer: ‘. . . I find it hard to understand that someone who arrives at such a conclusion does not seek the error in his argument’ (vKlOl). For vK ‘. . . the mind of the observer is irrelevant . . . the quantum mechanical measurement is terminated when the outcome has been macroscopically recorded . . . ” (vKlOl). Moreover, for vK, no special dynamics comes into play at ‘measurement’: ‘. . . The measuring act is fully described by the Schrödinger equation for object system and apparatus together. The collapse of the wavefunction is a consequence rather than an additional postulate . . . ” (vK97).

After the measurement the measuring instrument, according to the Schrödinger equation, will admittedly be in a superposition of different readings. For example, Schrödinger’s cat will be in a superposition |cat> = a|life> + b|death>. And it might seem that we do have to deal with ‘and’ rather than ‘or’ here, because of interference: ‘. . . for instance the temperature of the cat . . . the expectation value of such a quantity G . . . is not a statistical average of the values Gll and Gdd with probabilities |a|2 and |b|2, but contains cross terms between life and death . . . ” (vK103).

But vK is not impressed: ‘The answer to this paradox is again that the cat is macroscopic. Life and death are macrostates containing an enormous number of eigenstates |l> and |d> . . .

|cat>=∑lal|l>+∑dbd|d>

. . . the cross terms in the expression for <G> . . . as there is such a wealth of terms, all with different phases and magnitudes, they mutually cancel and their sum practically vanishes. This is the way in which the typical quantum mechanical interference becomes inoperative between macrostates . . .’ (vK103).

This argument for no interference is not, it seems to me, by itself immediately convincing. Surely it would be possible to find a sum of very many terms, with different amplitudes and phases, which is not zero? However, I am convinced anyway that interference between macroscopically different states is very, very elusive. Granting this, let me try to say what I think the argument to be, for the collapse as a ‘consequence’ rather than an additional postulate.

The world is again divided into ‘system’, ‘apparatus’, and the rest: W = S + A + R‘ = S‘ + R‘. At first, the usual rules for quantum ‘measurements’ are assumed at the S’/R’ interface – including the collapse postulate, which dictates correlations between results of ‘measurements’ made at different times. But the ‘measurements’ at S’/R’ which can actually be done, FAPP, do not show interference between macroscopically different states of S‘. It is as if the ‘and’ in the superposition had already, before any such measurements, been replaced by ‘or’. So the ‘and’ has already been replaced by ‘or’. It is as if it were so . . . so it is so.

This may be good FAPP logic. If we are more pedantic, it seems to me that we do not have here the proof of a theorem, but a change of the theory – at a strategically well chosen point. The change is from a theory which speaks only of the results of external interventions on the quantum system, S‘ in this discussion, to one in which that system is attributed intrinsic properties – deadness or aliveness in the case of cats. The point is strategically well chosen in that the predictions for results of ‘measurements’ across S’/R’ will still be the same . . . FAPP.

Whether by theorem or by assumption, we end up with a theory like that of LL, in which superpositions of macroscopically different states decay somehow into one of the members. We can ask as before just how and how often it happens. If we really had a theorem, the answers to these questions would be calculable. But the only possibility of calculation in schemes like those of KG and vK involves shifting further the shifty split – and the questions with it.

For most of the paper, vK’s world seems to be the petty world of the laboratory, even one that is not treated very realistically: ‘. . . in this connection the measurement is always taken to be instantaneous . . .’ (vKlOO)

But almost at the last moment a startling new vista opens up – an altogether more vast one:

‘Theorem IX: The total system is described throughout by the wave vector Ψ and has therefore zero entropy at all times . . .

This ought to put an end to speculations about measurements being responsible for increasing the entropy of the universe. (It won’t of course.)’ (vKlll)

So vK, unlike many other very practical physicists, seems willing to consider the universe as a whole. His universe, or at any rate some ‘total system’, has a wavefunction, and that wavefunction satisfies a linear Schrödinger equation. It is clear, however, that this wavefunction cannot be the whole story of vK’s totality. For it is clear that he expects the experiments in his laboratories to give definite results, and his cats to be dead or alive. He believes then in variables X which identify the realities, in a way which the wavefunction, without collapse, can not. His complete kinematics is then of the de Broglie-Bohm ‘hidden variable’ dual type: Ψ(t,q), X(t)).

For the dynamics, he has exactly the Schrödinger equation for Ψ, but I do not know exactly what he has in mind for the X, which for him would be restricted to some ‘macroscopic’ level. Perhaps indeed he would prefer to remain somewhat vague about this, for

‘Theorem IV: Whoever endows Ψ with more meaning than is needed for computing observable phenomena is responsible for the consequences . . .’ (vK99)

Towards a precise quantum mechanics

In the beginning, Schrödinger tried to interpret his wavefunction as giving somehow the density of the stuff of which the world is made. He tried to think of an electron as represented by a wavepacket – a wavefunction appreciably different from zero only over a small region in space. The extension of that region he thought of as the actual size of the electron – his electron was a bit fuzzy. At first he thought that small wavepackets, evolving according to the Schrödinger equation, would remain small. But that was wrong. Wavepackets diffuse, and with the passage of time become indefinitely extended, according to the Schrödinger equation. But however far the wavefunction has extended, the reaction of a detector to an electron remains spotty. So Schrödinger’s ‘realistic’ interpretation of his wavefunction did not survive.

Then came the Born interpretation. The wavefunction gives not the density of stuff, but gives rather (on squaring its modulus) the density of probability. Probability of what, exactly? Not of the electron being there, but of the electron being found there, if its position is ‘measured’.

Why this aversion to ‘being’ and insistence on ‘finding’? The founding fathers were unable to form a clear picture of things on the remote atomic scale. They became very aware of the intervening apparatus, and of the need for a ‘classical’ base from which to intervene on the quantum system. And so the shifty split.

The kinematics of the world, in this orthodox picture, is given by a wavefunction (maybe more than one?) for the quantum part, and classical variables – variables which have values – for the classical part: (Ψ(t,q . . .), X(t)… … ). The Xs are somehow macroscopic. This is not spelled out very explicitly. The dynamics is not very precisely formulated either. It includes a Schrödinger equation for the quantum part, and some sort of classical mechanics for the classical part, and ‘collapse’ recipes for their interaction.

It seems to me that the only hope of precision with the dual (Ψ, x) kinematics is to omit completely the shifty split, and let both Ψ and x refer to the world as a whole. Then the xs must not be confined to some vague macroscopic scale, but must extend to all scales. In the picture of de Broglie and Bohm, every particle is attributed a position x(t). Then instrument pointers – assemblies of particles have positions, and experiments have results. The dynamics is given by the world Schrödinger equation plus precise ‘guiding’ equations prescribing how the x(t)s move under the influence of Ψ. Particles are not attributed angular momenta, energies, etc, but only positions as functions of time. Peculiar ‘measurement’ results for angular momenta, energies, and so on, emerge as pointer positions in appropriate experimental setups. Considerations of the KG and vK type, on the absence (FAPP) of macroscopic interference, take their place here, and an important one, in showing how usually we do not have (FAPP) to pay attention to the whole world, but only to some subsystem and can simplify the wavefunction . . . FAPP.

The Born-type kinematics (Ψ, X) has a duality that the original ‘density of stuff picture of Schrödinger did not. The position of the particle there was just a feature of the wavepacket, not something in addition. The Landau-Lifshitz approach can be seen as maintaining this simple nondual kinematics, but with the wavefunction compact on a macroscopic rather than microscopic scale. We know, they seem to say, that macroscopic pointers have definite positions. And we think there is nothing but the wavefunction. So the wavefunction must be narrow as regards macroscopic variables. The Schrödinger equation does not preserve such narrowness (as Schrödinger himself dramatized with his cat). So there must be some kind of ‘collapse’ going on in addition, to enforce macroscopic narrowness. In the same way, if we had modified Schrödinger’s evolution somehow we might have prevented the spreading of his wavepacket electrons. But actually the idea that an electron in a ground-state hydrogen atom is as big as the atom (which is then perfectly spherical) is perfectly tolerable – and maybe even attractive. The idea that a macroscopic pointer can point simultaneously in different directions, or that a cat can have several of its nine lives at the same time, is harder to swallow. And if we have no extra variables X to express macroscopic definiteness, the wavefunction itself must be narrow in macroscopic directions in the configuration space. This the Landau-Lifshitz collapse brings about. It does so in a rather vague way, at rather vaguely specified times.

In the Ghiradi-Rimini-Weber scheme (see the box and the contributions of Ghiradi, Rimini, Weber, Pearle, Gisin and Diosi presented at 62 Years of Uncertainty, Erice, 5-14 August 1989) this vagueness is replaced by mathematical precision. The Schrödinger wavefunction even for a single particle, is supposed to be unstable, with a prescribed mean life per particle, against spontaneous collapse of a prescribed form. The lifetime and collapsed extension are such that departures of the Schrödinger equation show up very rarely and very weakly in few-particle systems. But in macroscopic systems, as a consequence of the prescribed equations, pointers very rapidly point, and cats are very quickly killed or spared.

The orthodox approaches, whether the authors think they have made derivations or assumptions, are just fine FAPP – when used with the good taste and discretion picked up from exposure to good examples. At least two roads are open from there towards a precise theory, it seems to me. Both eliminate the shifty split. The de Broglie-Bohm-type theories retain, exactly, the linear wave equation, and so necessarily add complementary variables to express the non-waviness of the world on the macroscopic scale. The GRW-type theories have nothing in their kinematics but the wavefunction. It gives the density (in a multidimensional configuration space!) of stuff. To account for the narrowness of that stuff in macroscopic dimensions, the linear Schrödinger equation has to be modified, in the GRW picture by a mathematically prescribed spontaneous collapse mechanism.

The big question, in my opinion, is which, if either, of these two precise pictures can be redeveloped in a Lorentz invariant way.

‘. . . All historical experience confirms that men might not achieve the possible if they had not, time and time again, reached out for the impossible.’ Max Weber

‘. . . we do not know where we are stupid until we stick our necks out.’ R P Feynman

The Ghiradi-Rimini-Weber scheme

The GRW scheme represents a proposal aimed to overcome the difficulties of quantum mechanics discussed by John Bell in this article. The GRW model is based on the acceptance of the fact that the Schrödinger dynamics, governing the evolution of the wavefunction, has to be modified by the inclusion of stochastic and nonlinear effects. Obviously these modifications must leave practically unaltered all standard quantum predictions about microsystems.

To be more specific, the GRW theory admits that the wavefunction, besides evolving through the standard Hamiltonian dynamics, is subjected, at random times, to spontaneous processes corresponding to localisations in space of the microconstituents of any physical system. The mean frequency of the localisations is extremely small, and the localisation width is large on an atomic scale. As a consequence no prediction of standard quantum formalism for microsystems is changed in any appreciable way.

The merit of the model is in the fact that the localisation mechanism is such that its frequency increases as the number of constituents of a composite system increases. In the case of a macroscopic object (containing an Avogadro number of constituents) linear superpositions of states describing pointers ‘pointing simultaneously in different directions’ are dynamically suppressed in extremely short times. As stated by John Bell, in GRW ‘Schrödinger’s cat is not both dead and alive for more than a split second’.

  • The original and technically detailed presentation of GRW can be found in 1986 Rev D 34 470; a brilliant and simple presentation has been given by John Bell in Schrödinger: Centenary Celebration of a Polymath C W Kilmister (ed) 1987 Cambridge University Press p41
  • A general discussion of the conceptual implications of the scheme can be found in 1988 Foundation of Physics 18 1
  • The GRW model has been the object of many recent papers and a lively debate on its implications is going on. Recently the model has been generalised to cover the case of systems of identical particles and to meet the requirements of relativistic invariance.

G C Ghiradi, A Rimini and T Weber

Further reading

J S Bell and M Nauenberg 1966 The moral aspect of quantum mechanics in Preludes in theoretical physics (in honour of V F Weisskopf) (North-Holland) 278-286

P A M Dirac 1929 Proc. R. Soc. A 123 714

P A M Dirac 1948 Quantum mechanics third edn (Oxford University Press)

P A M Dirac 1963 Sci. American 208 May 45

K Gottfried 1966 Quantum mechanics (Benjamin)

K Gottfried Does quantum mechanics describe the collapse of the wavefunction? Presented at 62 Years of Uncertainty, Erice, 5-14 August 1989

L D Landau and E M Lifshitz 1977 Quantum mechanics third edn (Pergamon)

N G van Kampen 1988 Ten theorems about quantum mechanical measurements Physica A 153 97–113

J von Neumann 1955 Mathematical foundations of quantum mechanics (Princeton University Press)

  • This article is published with the permission of Plenum Publishing, New York; it appeared in the proceedings of 62 Years of Uncertainty (Erice, 5–14 August 1989)

Shifts in atomic understanding

Technology is set to yield, in the near future, unprecedentedly short (10–14 s) and intense (1017–1019 W cm–2) light pulses. Such intensities have in the past been the preserve of the gigantic systems associated with laser fusion research but a considerable reduction in system size has made them available to the atomic physicist. Experiments using such lasers are just beginning and there have been as yet few investigations at intensities beyond 1014 W cm–2. Nevertheless these have already provided some delightful surprises.

One is that, contrary to general belief, atomic structure does still play an important role at these intensities – thus such pulses provide a powerful tool to investigate atomic distortions in intensity ranges where standard techniques fail. In a sense, ultrashort pulses allow us to take snapshot views of the atoms within the laser field. For instance, recent experiments on multiphoton ionisation have revealed giant energy shifts in atoms subjected to 5 × 1013 W cm–2.

The atomic energy level structure is governed, of course, by the interaction between the electron and the nucleus. For the simplest atom, hydrogen, for instance, this structure is very simple: in the non-relativistic limit, the energy levels depend only on the principal quantum number n, as (–1/2n2), and occur closer and closer to each other as n goes to infinity. In this case the energy does not depend on angular momentum, states with the same principal quantum number and different angular momenta being degenerate. At low intensities the light induces transitions between unperturbed atomic states.

At high intensities, in contrast, the interaction between the electron and the nucleus may be dominated by the interaction between the electron and the light field. The energy level structure is therefore, in general, strongly distorted and the degeneracy is suppressed. In any case, the electron interaction with the laser light can no longer be considered as a small perturbation of the atomic system. The system, atom and light, must be treated as a whole. Thus this physical domain is often characterised as nonperturbative. It is very likely that contiguous disciplines such as plasma physics or surface physics will soon benefit from developments in this regime, a prospect that brightens the outlook of many a research group.

Multiphoton physics

Elementary quantum mechanics tells us that transitions between two bound states give rise to a discrete absorption/emission spectrum, whereas transitions to unbound or free states (ionisation) yield a continuous spectrum. Energy conservation in ionisation determines the kinetic energy of the outgoing electron Ek = hv – E1 where hv is the energy of the photon of frequency v and E1 is the binding energy of the initial atomic state. Another class of processes is formed by the transitions involving the simultaneous absorption of several photons. Such multiphoton transitions allow one, in principle, to ionise an atom with light corresponding to photon energies smaller than E1 provided photons N are absorbed simultaneously. Energy conservation then yields an electron with kinetic energy Ek = Nhv – E1.

Multiphoton ionisation (MPI) processes are observable only with intense pulsed laser sources once N is larger than two. If the intensity is not too high, MPI processes of order greater than two can reasonably be described within the framework of time-dependent perturbation theory, to the lowest ‘non-vanishing’ order. This means that, if the absorption of a minimum number of photons N is required, for energy conservation, to ionise the atom, perturbative solutions of the Schrodinger equation up to order N must be found to obtain the ionisation rate. Such an approximation is valid if the shifts and broadenings of atomic states can be neglected, or at least for all ‘non-resonant’ transitions observed with intensities up to 1011–1012 W cm–2. By ‘non-resonant’ we mean transitions in which the atom ground state is coupled to the continuum without going through some intermediate excited or ‘resonant’ state. Under this approximation, the ionisation rate scales as the laser intensity to the power N.

In fact, part of this description breaks down as soon as the intensity I is around 1011 W cm–2 as immediately revealed by electron energy analysis (figure 1a). Atoms can absorb not only the N photons necessary for ionisation but also a few more (S) which are used to accelerate the electrons, increasing their kinetic energy by Shv. The electron energy spectra display a number of lines of small and decreasing amplitude, separated by the photon energy. The production rate of each ‘peak’ scales as IN+1, IN+2 . . . This process, known as above-threshold ionisation (ATI), is a correction to MPI of the lowest possible order and was first detected at Saclay.

Moving to high intensities

Things become more complicated as the intensity is increased (>1012 W cm–2) because it is not then possible to assume that transitions occur between unperturbed states. Suppose one electron is orbiting around the nucleus. When the electromagnetic field (produced by the laser) is turned on, the Lorentz force perturbs the electron motion and therefore its energy. Changes in the energy of atomic bound states then induced are known as AC-Stark effects because of their relation to the Stark effect.

Three situations might occur: (i) if the light or field frequency v is close to the energy difference between the initial state and some state to which it can be coupled, the two states are ‘mixed’ and their energies are shifted by this mixing. If v is smaller than their energy difference then the two states are pushed apart, and vice versa if v is larger than their energy difference. (ii) If v is small compared to any atomic energy difference then the state the electron occupies is repelled by the coupling to the other states. This is usually the case for the atom ground state if v is in the red or infrared region of the electromagnetic spectrum, and the ground state is down-shifted in energy by a small amount. (iii) If v is large compared to any atomic energy difference (for instance for states whose energies lie close to the ionisation threshold) then the shift is always positive and large. In this case there is a simple classical interpretation of the shift: the motion of the electron can be thought of as a slow orbiting motion around the nucleus superposed by a fast ‘quiver’ motion driven by the electromagnetic field around the average orbit. The latter has an average kinetic energy which adds to the electron energy. Its value in eV is easily computed from classical mechanics, given by EQ = 0.94 × 10–13 I λ2 where I is the intensity in W cm–2 and λ the wavelength in μm. This energy is known in the literature as the quiver energy or the ponderomotive potential. The same interpretation does not however hold for strongly bound electrons since the orbital motion is then faster than the oscillatory motion.

Figure 1

graph

(a) Photoelectron energy spectrum from multiphoton ionisation of xenon (ionisation potential 11.13 eV) by 1.17 eV photons (intensity = 2 × 1012 W cm–2). The leftmost peak is produced by 11 photon ionisation (the minimum required by energy conservation). The next two peaks, due to 12 and 13 photon ionisation, are separated by an energy equal to the photon energy (above threshold ionisation); (b) As in (a) but with a five times larger intensity (1013 W cm–2). Ten photons are absorbed above the ionisation threshold. The strong electromagnetic field has distorted the atomic structure; the first peak is totally suppressed while the most probable one is the fourth one.

Stark shifts as described above become crucial once the intensity is above 1012 W cm–2 and have two consequences: first, the small down-shift of the ground state and the large up-shifts by the quiver energy of the states close to the continuum causes the ionisation potential to be apparently increased by a little more than the quiver energy. For intensities of 1013 W cm–2 and a wavelength of 1 μm this increase is about 1 eV, that is, the energy of one photon. Therefore ionisation of the atom requires absorption of at least (N+1) photons. This leads to spectra such as those in figure 1b: the first peak which would correspond to N photon ionisation is missing; the second one which corresponds to (N+1) photon ionisation is strongly suppressed while the most probable peak corresponds to an (N+2) photon transition.

A second consequence is that the energy of some state of the atom may be shifted via the Stark effect in such a way that it becomes equal at some intensity to the energy of an integer number of photons. The transition probability is then boosted if this resonance condition is met. For instance if one keeps the intensity constant and scans the light frequency, the ionisation signal is strongly enhanced in a small region of frequency; if the frequency is kept constant and the intensity scanned, then a similar enhancement is seen around the intensity value which Stark-shifts the state into resonance. In a real focused laser beam, however, this is unnecessary as all intensities between zero and the maximum are already present. If, in some part of the focused region, the intensity causes resonance, then more electrons will be produced there because the probability is much higher. The electrons produced in this region of the beam have an energy Ek = (N+S)hv – (E1+EQ(IR)) where IR is the resonant intensity: neglecting the small shift of the ground state, the ionization potential E1 appears to be increased by EQ(IR). As a consequence, the electron energy spectrum should show a bump at the energy Ek, which is shifted by EQ from (N+S)hv – E1. But this is not what is shown in figure 1b. No extra bumps are detected and the peaks are exactly at (N+S)hv – E1, although the intensity is high enough for the first peak to be suppressed, and shifts of the order of the separation between two adjacent peaks should be observed.

After the first observation at the FOM Institute in Amsterdam, it took some time to understand that this discrepancy occurs because the electron energy was measured outside the laser beam. This apparently innocent circumstance has drastic consequences for the measured energy. The best way to see this is the following: the photoelectrons are created inside the laser beam with both a translation energy Ek and a quiver energy EQ (equivalent to saying that the ionisation potential is increased by EQ). As an electron travels to the detector which lies outside the beam, it has to cross the region where the intensity is changing rapidly, from the peak intensity to zero over typically 10 μm. In doing so it is subjected to the Lorentz force which depends both on time and space. Classical electrodynamics shows that this force has a component proportional to the gradient of the electromagnetic field. This ‘ponderomotive force’ tends to accelerate the electron towards the low intensity regions, that is outside the focused beam. Furthermore, a classical calculation shows that the total work done by this force along the electron trajectory is precisely EQ. In fact, the quiver energy acts as a potential (hence its other name, the ponderomotive potential) from which the ponderomotive force is derived. Again from classical mechanics, the change in kinetic energy is shown to be equal to the variation of the potential, which is exactly EQ since the quiver energy is zero outside the beam. Hence, the electron energy outside the beam is just (N+S)hv – E1 no matter what value the initial quiver energy of the photoelectron has (provided S is sufficiently large, that is for the peaks which are not suppressed).

Short sharp pulses

This is where ultrashort pulses are really useful. If the pulse duration is shorter than the time taken by the electron to leave the beam (which depends on the initial electron velocity and the size of the beam) then some or all of the quiver energy is lost. Changing the laser pulse from 136 ps to 50 ps was sufficient to allow us to observe the predicted shifts in the energy peaks (see figure 2). The slow electrons (low energy) are shifted more than the fast ones (high energy) and the shift depends on the ratio of the pulse duration to the electron exit time. With this observation, it was easy to figure out what would happen if, instead of a 50 ps pulse, a much shorter pulse (1 ps or less) was used: the pulse would be much shorter than all exit times and all the peaks would be shifted by the same amount, EQ.

Figure 2

graph

Photoelectrons are produced with both translation energy and quiver energy. (a) Long-pulse (136 ps), low-intensity reference spectrum; (b) In long pulses (136 ps) the quiver energy is converted into translation energy as the electron travels out of the laser beam and the peaks are not shifted even though some may be suppressed; (c) For short pulses (50 ps), part of the quiver energy is lost and the slower the electrons the more the peaks are shifted towards lower energies. Vertical solid lines show the expected positions of the peaks.

The technique of generating subpicosecond pulses is now well known: they are produced from CW, mode-locked ring cavities working in the colliding pulse, mode-locked mode. In such a laser, the gain is provided by Rhodamine dye, pumped by a CW argon ion laser, while the mode-locking is achieved by a saturable absorber. Two pulses counter-propagate within the ring cavity and collide inside the mode-locking dye jet, hence the name. Extremely short pulses (30 fs) with peak power of the order of a few kilowatts are produced at a repetition rate equal to the round-trip time of the pulse in the cavity. To reach the required intensities of 1014 W cm–2 the pulses are further amplified by YAG pumped amplifiers.

When subpicosecond pulses were used, the expected shifts, equal for all ATI peaks, were seen. Furthermore, the electron energy spectra displayed a number of extra substructures due (as we now know) to Stark-tuned resonances as explained above. When in fact this was first observed by a research group at Bell Labs, the explanation was far from obvious. From studies of resonant MPI by the usual technique of tuning the laser frequency, it was known that, as the intensity is increased, the resonances get broader and broader and, eventually, become undetectable. It was therefore generally taken for granted that atomic structure did not play any important role at high intensity. Of course, it took some more work to prove that the extra bumps, which looked like ‘noise’ in the first spectra, were due to atomic resonances.

Figure 3

graph

(a) Photoelectron energy spectrum from multiphoton ionisation of xenon by 2 eV photons produced by a linearly polarised femtosecond laser: the main structure of three peaks separated by the energy of one photon is subdivided into more subpeaks that reveal the Stark induced resonances; (b) As in (a) but with circularly polarised light. Selection rules forbid the resonances, leaving only the main structure of the spectrum. This proves the atomic origin of the substructure in (a).

One crucial test was carried out by using circularly polarised light, in which all the photons have the same angular momentum, say +1, contrary to linearly polarised light which contains as many photons with an angular momentum +1 as photons with angular momentum –1. As a consequence, in an N photon transition with circularly polarised light, the atom which absorbs one unit of angular momentum with each absorbed photon, must increase its total angular momentum by N units. For instance, a six photon transition starting from an S state (l=0) must end in an I state (l=6) to conserve the total angular momentum. With circularly polarised light, all the extra substructures in the electron energy spectra were suppressed (figure 3b), leaving only the normal ATI ‘comb’. This not only confirmed the atomic origin of the resonances but also gave valuable information on the angular momenta of the resonant states. This and other tests led to the important conclusion that all multiphoton ionisation processes are resonant, whatever the atom and the laser frequency. Because of the combined effects of the shifts due to the loss of the quiver energy and the increase of probability due to the Stark-tuned resonances, the electron energy spectra taken with ultrashort pulses are an image of the ionisation probability as a function of intensity: such a function is clearly no longer a simple power law but a very complicated succession of maxima due to the resonances. This role of the atomic structure could only have been unveiled by intense ultrashort pulses and would have remained hidden in experiments using long pulses. Ultrashort pulses give a snapshot of the atomic energy levels as they are inside the light beam.

Of course, the larger the detuning from the energy of a given state and the energy of an integer number of photons, the higher the intensity needed to Stark-tune the state to resonance. Consider a small volume inside the focused laser beam where such an intensity is reached during the pulse. This could be thought to happen twice – once during the rise time and once during the fall time of the pulse – but, in fact, careful analysis shows that only the atoms in the region where the required intensity for a given resonance to occur is achieved around the maximum of the pulse will contribute to that resonance. Whatever the detailed mechanism, it appears as though the atom is itself picking the right intensity. The drawback is that substructures will occur at the same energy in the photoelectron spectra whatever the pulse peak intensity. Therefore, even though the peak intensity was varied to get information on the distorted atomic structure only one answer would come out of the type of experiments described above.

Figure 4

graph

Stark-induced resonances. Energy levels are shown for three different electromagnetic field intensities: (a) at very weak intensity, the levels are not shifted and three-photon ionisation is energetically allowed. The dotted arrow indicates the absorption of one extra photon (ATI); (b) the intensity is sufficient to shift one of the excited states (ES) into resonance with the energy of two photons. The ground state (GS) and the ionisation threshold (IT) are shifted and the three-photon ionisation becomes energetically forbidden; (c) at even higher intensity the other excited state is shifted into resonance and the ionisation potential increases further; (d) the ionisation rate as a function of intensity: each time the intensity adjusts a resonance, the rate increases and then decreases as the intensity shifts the level out of resonance.

The way out of this dead end is to tune the laser. For a given resonance this changes the detuning and hence the intensity needed to Stark-tune the state into resonance. Following the position of the corresponding substructure as a function of the laser frequency will yield the energy of the state as a function of intensity, that is the dynamics of the Stark shifts at high intensity. Making the CPM laser tunable is very difficult, but fortunately, there is a trick to make life easier. If the output of the amplified femtosecond pulse is focused into a simple cell containing water then a white light continuum is generated through self-phase modulation. A narrow part of this continuum can be selected in an interference filter and subsequently amplified by another series of YAG-pumped dye cells. The experiment outlined above may be carried out for seven photon ionisation of xenon. By tuning the frequency as described, it was possible to change the detuning of six-photon resonances (6hv – ER, where ER is the unperturbed energy of a state which may be pushed into resonance by the Stark shift) by several electron volts. The electron volt is an energy unit which is really enormous by spectroscopic standards: it is equal to 8066 cm–1, the cm–1 being the standard unit in spectroscopy. Even in atomic units the electron volt is a huge quantity since the ionisation potential (half the atomic unit of energy) of hydrogen, for instance, is 13.6 eV. For a state which shifts as a free electron, an intensity of several 1013 W cm–2 is needed to push it upwards by 1 eV.

In the range explored in the experiment, the energies of all identified states depended linearly on intensity with a coefficient very close to the one found at very low intensity. This was very surprising since such a linear dependence was expected to break down rapidly as the intensity was increased. The experiment seems to show that atomic states dressed by a very intense electromagnetic field have a ‘simple’ behaviour although the reason for this is not yet fully understood.

Recent non-perturbative calculations show that there are three main intensity regimes: (i) very low – the shifts are predicted by second order perturbation and are linear in intensity; (ii) higher – the shifts are no longer linear and the states are strongly mixed energy levels; (iii) high – new states emerge which are strongly mixed too, even if they can be followed by continuity with a ‘naked’ state. Calculations suggest that these states may sometimes shift again linearly with intensity as in the first regime.

Obviously much more work is necessary to reach quantitative comparisons between theory and experiment. But, for the first time, it appears possible to study atoms dressed by very strong photon fields.

Two features round off this picture of MPI in ultrashort pulses. First there is the suppression of resonances under certain conditions. It is known that, by simultaneously exciting a number of atomic states close to the ionisation limit (‘Rydberg’ states) a wavepacket is created which behaves more or less as a classical electron – the wavepacket is concentrated in space and distance to nucleus varies periodically with time. The probability that such a state will absorb photons varies periodically with time also and is a maximum when the wavepacket is close to the nucleus. When the ‘wavepacket’ is far from the nucleus it behaves essentially as a ‘free’ electron and a free electron cannot absorb photons because momentum cannot be conserved. If the pulse that excites the Rydberg states is so short that it turns off before the wavepacket returns from being close to the nucleus, then there is no time to absorb more photons while the atom is in the excited state and the resonance, which would occur in a long pulse, is suppressed.

Secondly, other experiments, carried out under rather different conditions, tend to show that the ionisation rate is much lower than predicted. The reason for this is not yet understood. For short pulses (less than 10 fs), the laser bandwidth becomes so large that many atomic states are resonant. Strangely enough, numerical simulations show that the ionisation rate is then decreased as if the electron remained trapped in the discrete spectrum.

New regions of research are opening up with the new tabletop ultra-powerful lasers. Intensities between 1019 and 1021 W cm–2 are likely to be achieved within the next year or so. A totally new class of processes is expected, ranging from the complete stripping of all electrons from atoms to vacuum polarisation phenomena: the generation of light at a frequency three times that of the light used – a very common process in dense media which at sufficient intensities becomes possible in a vacuum. The results summarised here may have been obtained at a more modest intensity, but they contain enough surprises to make one think that many more are waiting for us in the very near future.

Further reading

P Agostini and G Petite 1988 Photoelectric effect under strong irradiation Contemp. Phys. 29 57

May 1987 Special issue of J. Opt. Soc. America B 4 765

1988 The ponderomotive potential of high intensity light and its role in multiphoton ionization of atoms IEEE J. Quantum Electron. 24 1461

Copyright © 2026 by IOP Publishing Ltd and individual contributors