Showing posts with label SCW. Show all posts
Showing posts with label SCW. Show all posts

08 November 2013

Too big for its boots

In a delightful way, from a writing perspective, my last few data analysis topics have first synergised with one another and then led naturally on to the consideration of Big Data.
What exactly is "big data"? The answer, it should come as no particular surprise to hear, is "it depends". As a broad, rough and ready definition, it means data in sufficient volume, complexity and velocity to present practical problems in storage, management, curation and analysis within a reasonable time scale. In other words, data which becomes, or at least threatens to become in a specific context, too dense, too rapidly acquired and too various to handle. Clearly, specific contexts will vary from case to case and over time (technology continuously upgrades our ability to manage data as well as generating it in greater volume) but broadly speaking the gap remains – and seems likely to remain in the immediate future. The poster boys and girls of big data, in this respect, are the likes of genomics, social research, astronomy and the Large Hadron Collider (LHC) whose unmanaged gross sensor output would be around fifty zettabytes per day.
There are other thorny issues besides the technicalities of computing. Some of them concern research ethics: to what extent, for example, is it justifiable to use big data gathered for other purposes (for example, from health, telecommunications, credit card usage or social networking) in ways to which the subjects did not give consent? Janet Currie (to mention only one recent example amongst many) suggests a stark tightrope with her "Big data vs big brother" consideration of large scale pædiatric studies. Others are more of concern to statisticians like me: there is a tendency for the sheer density of data available to obscure the idea of a representative sample– and a billion unbalanced data points can actually give much less reliable results than thirty well selected ones.
Conversely, however, big data can also be defined in terms not of problems but of opportunity. Big data approaches open up the opportunity to explore very small but crucial effects. They can be used to validate (or otherwise) smaller and more focussed data collection, as for instance in Ansolabehere and Hersh’s study [1] of survey misreporting. As technology gives us expanding data capture capabilities at ever finer levels of resolution, all areas of scientific endeavour are becoming increasingly data intensive. That means (in principle, at least) knowing the nature of our studies in greater detail than statisticians of my generation could ever have dreamed. A while back, to look at the smaller end of the scale, I mentioned [2] the example of an automated entomological field study régime simultaneously sampling two thousand variables at a resolution of several hundred cases per second. That’s not, by any stretch of the imagination, in LHC territory but it’s big enough data to make significant call on a one terabyte portable hard drive. It’s also a goldmine opportunity for small team or even individual study of phenomena which would not long ago have been beyond the reach of even the largest government funded programme: big data has revolutionised small science.
There is, in any case, no going back; big data is here to stay – and to grow ever bigger, because it can. Like all progress, it’s a double edged sword and the trick as always is to manage the obstacles in ways which deliver the prize. [more]

[1] Ansolabehere, S. and E. Hersh, Validation: "What Big Data Reveal About Survey Misreporting and the Real Electorate". Political Analysis, 2012. 20(4): p. 437-459.

[2] Grant, F., "Retrieving data day queries", in Scientific Computing World. 2013, Europa Science: Cambridge. p. 10-12..

09 August 2013

Pretty as a picture

Talking about statistical work by nonstatisticians, recently ("Stats for the million", 14 June), I mentioned the importance in that context of graphical visualisation of data. It goes well beyond that, however.
On the one hand, fuelled by the ever-accelerating growth curve in computing power per unit of investment, visualisation has progressively moved to the core of exploratory and analytic strategies. The effects on traditional methods are profound, as separate work phases collapse into continuous cybernetic feedback loops and statisticians develop increasingly immersive relationships with their raw material. On the other, data visualisation has penetrated mainstream discourse to become an integral part of vernacular literacy – “one of the genuinely new cultural forms enabled by computing” as Lev Manovich [1, 2] describes it.
Those two aspects, the technical and the vernacular, are not separate; they are two sides of the same coin. They are beginning to interpenetrate with other developments such as direct onscreen haptic manipulation of program interfaces and may in the long run turn out to be the most far reaching and profound effect of the scientific computing revolution.
At the heart of this lies the capacity of inexpensive desktop, laptop or even handheld devices to manipulate graphics in real time response to user curiosity. When I started writing for Scientific Computing World, back in the 1990s, it was possible to represent three data variables as a scatter plot cloud, or as a fitted surface, on x, y and z axes, but changing the viewpoint or scale usually involved typing new parameters into a settings box and watching the screen progressively redraw. It seemed pretty cool, then. I remember my excitement when the major statistics packages, one by one, added the ability to grab the plot with a mouse click and intuitively apply zoom, pitch, roll and yaw by dragging. Nowadays, I can do the same on a pocket tablet or even a cellphone by simply sliding my fingers around the image itself. On a desktop, laptop or heavier tablet machine I have access to considerably more than three dimensions, not to mention different display types such as vector flows in the same visualisation as positional points, planes or volumes.
Not that such impressive psychoperceptual pyrotechnics are always necessary or even desirable in every context. Detailed 2D presentation of very traditional plots of the kind that would have been familiar to my primary school self in the late 1950s are, in many circumstances, still the best visualisations of real world situations. The miracle of current software is that those two extremes, and everything between, are available off the shelf to suit the needs of the moment. [more...]

  1. Lev Manovich, “Data visualisation as new abstraction and anti-sublime” in Small tech: the culture of digital tools, electronic mediations, B. Hawk and D.M. Rieder, Editors. 2008, University of Minnesota Press: Minneapolis.
  2. Lev Manovich, Software takes command : extending the language of new media. International texts in critical media aesthetics.  9781623568177.
Full references list here

14 June 2013

Stats for the million

Most of my consultancy work boils down to devising ways in which clients can maximise the validity and quality of work by lay staff or volunteers whilst minimising the impact of inevitable errors. And my email inbox is rarely without a plaintive request (open or veiled) for informal guidance or reassurance on a conclusion in which the sender lacks the confidence necessary for exposing it to peer scrutiny.
In the real world, whether we statisticians like it or not, an overwhelming majority of statistically based decisions and statistically oriented actions are taken by nonstatisticians.
Having made such a sweeping statement, it would be prudent to take a moment for definition of terms. What, exactly, is a nonstatistician?
In my own practice, there is no single clear cut answer to that. The word is not the binary divisor which it pretends to be: just a shifting pointer on a continuum. That’s why I used the alternative description "lay staff" in my first paragraph, above: because it’s less complicated. All too often, individuals with the same level of expertise will place themselves very emphatically on opposite sides of the pointer. In one market segment with which I am very familiar, professionals who are hazy about the difference between mean and mode are counted as statisticians and make million euro decisions based on statistical grounds. In another, science graduates whose degree transcripts include passes in all the usual statistics courses claim to be completely baffled by the subject.
Reading through the literature of any professional area, but particularly in the medical and life sciences, one finds frequent reference to this. [more...]

05 April 2013

Retrieving data day queries

Perhaps the most famous data retrieval case in the history of science comes from sixteenth century orbital mechanics. Copernicus had laid the foundations for a viable heliocentric system; Kepler stood ready to finalise it. Between the two, both problem and solution, lay the mysteries of Mars: "the wanderer planet". The data which Kepler needed already existed, in a database of naked eye observations painstakingly constructed over two decades by Danish philosopher Tycho Brahe.
The problem was twofold. Brahe had nailed his colours to a mixed system at odds with that of Copernicus; and his data were his claim to posterity. He employed Kepler as an assistant, but jealously guarded access to the full observational data set.
Kepler did, eventually, gain access to the data. It wasn’t easy, nor always amicable (though allegations that he murdered Brahe to achieve it have been discounted), but it was done. He still had to learn how to retrieve it productively, but six years of mining and analysis finally bridged the gap to produce a final, successful, validated model.
Things have changed almost unrecognisably over the four or five centuries since Copernicus, Kepler and Brahe, but some features recognisably remain amid the new. Investment in research is balanced against the advantages of shared access. Boundaries, proprietary or otherwise, remain between researchers and data repositories. Murder and less extreme espionage methods may be rare (though not unheard of) as means of gaining access to data stores, but Kepler would no doubt recognise in essence the processes of negotiation and persuasion which allow those boundaries to be permeated.
The biggest early twenty first century data retrieval issue, however, is a different one. Acquisition in large quantity is becoming ever easier. Storage is, in relative terms, becoming cheaper. The headache often becomes how to ensure that one retrieves the right data for particular purposes from the ever ballooning volumes which are thus becoming available.
And then there is the problem of storage format obsolescence. Unlikely as it may seem, digital information which is by definition recent and (you might think) ought to be more easily accessible, and more carefully curated, can sometimes be harder to reach than older analogue stores. [More...]

21 February 2013

Omnilingual word spotting

In yesterday's out take from my recent “Joy of text” piece in Scientific Computing World, I mentioned in passing the term “word spotting”.
Word spotting looks for visually similar discrete components within a text, and classifies them using statistical comparisons. In approximate human terms, it is treating words (or combinations of words) as ideograms rather than as phonetic constructs.
Word spotting is not limited to written or printed material, though that is the context with which I'm concerned here: it also applies, for example, to speech recognition. Nor, intriguingly, is it necessarily limited to words of known meaning; it can equally well be applied to semantic units of entirely unknown signification. It could, as an extreme example, be used to analyse the manuscripts from a lost extraterrestrial civilisation in H Beam Piper's classic science fiction story Omnilingual.
Reverse the Omnilingual example to imagine a hypothetical extraterrestrial archaeologist trying to study post apocalyptic remains of our own cultures. It will be obvious from context that certain signification units are associated with the physical sciences: "volts" on electrical signs and appliances, just to pick one.
Our xenoarchaeologist (who may not have alphabetic scription systems, or even vocalisation, and certainly cannot assume that each letter represents a sound) looks at a mass of textual matter whose content, subject matter, purpose and reliability are unknowable. Where should attention be concentrated? The only certain knowledge is that words are identifiable visual entities found in isolation and occurring with spatial separation in books.
Word spotting, with no semantic assumptions, quickly shows that some books contain many instances of the visually related signifiers "volts" "volt", "voltage", "voltmeter" and so on, suggesting that those sections contain material related to electricity; others do not. There will, of course, be false hits such as "Voltaire" and "revolt", but as one discriminator amongst others in a multiple sieving process it would nevertheless be invaluable.
Handwritten notes and journals would be less amenable than printed books (in my own handwriting, for instance, computer transcription systems have trouble separating "volt" from "bolt" and sometimes even "void"), but could still be sieved using multiple discriminators in the same way.

  • Piper, H.B., Omnilingual, in Astounding Science Fiction. 1957, Dell Magazines: Northwalk CT.

20 February 2013

Making sense of the census

In preparing an article on any topic, there is often more good stuff than makes it into the final edit.
In the case of my recent “Joy of text” piece, about text analysis for Scientific Computing World, one of the out takes was a prototype content-based image retrieval system framework developed for census searches by Kenton McHenry and his ISDA group at NCSA. The problem to be solved is computerised searching of large volume handwritten census returns.
A user inputs a handwritten query – I might, for instance, write "Grant". The system derives a numerical feature vector which describes that input, then seeks occurrences of similar vectors within the image database.
The system is designed to self-validate, by recording which returned entities are selected by the user. I will not, for example, select false hits such as "Grand" or "Ghent" and the system will note which results I do or do not follow up. Over time, other Grants will make similar decisions and increase the system's confidence in selecting some hits for return and not others; gradually, those writing “Grant” as their query will see fewer and fewer offers of documents containing similar looking words.
The computer analytic process behind all this is progressive.
First, the lines and boxes on the census forms are used to carve up the content into image segments (for instance, surname will be in a box at the same location on each form and will become a data entity). Each segment is then converted into a numerical feature vector representing its appearance, and similar feature vectors are grouped hierarchically. Two million XSEDE (Extreme Science and Engineering Discovery Environment) CPU hours have been requested for initial record processing.
When the search query is entered, word spotting is used to compare its vector with those stored in the database, seeking matches within statistically defined limits of similarity. The search is not a blind one from the beginning of the database through seventy billion image segments to the end; the hierarchical grouping guides greatly reduces the number of entities which need to be compared.

15 February 2013

The joy of text

It's a mundane truism, not normally worth mentioning, that words and phrases as signification units in natural language have only the fuzziest of relations to that which they signify. It is, nevertheless, a live issue for the many researchers attempting to computerise data analytic activity using text as raw material. It's also a truism of which I have been reminded afresh as I discussed the topic with practitioners and consumers of textual analysis, no two of whom used the term in exactly the way.
Strictly speaking, textual analysis describes a social sciences methodology for examining and categorising communication content. In practice, though, it is widely used to cover a range of activities in which unstructured or partially structured textual material is submitted to rigorous analytic treatment. What they all have in common is a desire to wrestle the petabytes of potentially valuable information locked up in an ever inflating text reservoir (blogs, books, chat rooms, clinical notes, departmental minutes, emails, field journals, lab notebooks, patents, reports, specification sheets, web sites and a million other sources) into a form which is susceptible to useful, objective data analytic treatment. Temis, of whom more later, have on their website a headline which sums it up neatly: "Big data issue #1: a lot of content and no insights". Text mining, the consequent knowledge bases, and analysis of the results have become a major component of biomedical and pharmaceutical research.
For our purposes here, I have taken it to mean analysis whose purpose is to extract scientific value from texts, to examine those texts scientifically, or some combination of the two. [More]

20 January 2013

Ode to OCR

Preparing an article on text analysis for the upcoming issue of Scientific Computing World (watch this space) I've been revisiting the various ways of getting physical text into digital form.
Optical Character Recognition (OCR) is the workhorse of text transcription and, while we all grumble about its very real real short comings shortcomings, it really does a remarkably good job, most of the time, of rendering graphic images of printed fonts into digitised text for analysis.
Even at the lowliest manual level, and with all its admitted faults, OCR is a useful tool. A colleague and I recently had to add a six hundred and fifty page eighteenth century text to a digitised textbase for analysis. The rare and valuable paper original which we located was in a library, and could not be removed. Filing a request for the digitisation to be carried out would take weeks. With the consent of the library we used a smartphone, a ten year old copy of ABBYY FineReader 5 (now in release 11, and correspondingly more developed, as part of a software range for different text tasks) and a netbook. Even allowing for manual error correction, we had our validated data within three hours and the library added a copy to its own digital records. A similarly sized text already available as graphic only PDF was transferred more quickly still.

18 December 2012

Going places

As a child I lived, for a while, on open plains at opposite sides of the planet where (like all my schoolmates) I was fascinated by glimpses of how transport might be in our future. The P1127 (an experimental VSTOL aircraft which would later become the Harrier); the Belvedere (a twin rotor helicopter, already ageing and soon to be swept away by the Chinook); and, most exciting of all, something which we heard but never saw.
This gave rise to one of my earliest serious data gathering and analysis projects: a notebook in which I religiously recorded details of everything mechanical that moved and, with particular attention to detail, rocket test firings. The date, time and direction for every roaring launch, every thud of a concrete warhead returned to earth, went into my little book. After a while, though I didn’t yet know the term “data analysis”, I was able to tell my friends with fair accuracy when the next launch was likely to be, how long the flight would take and where the impact would probably be.
These rockets had mysterious names like Honest John, Thor, Thunderbird, and Bloodhound. Fifty-or-so years on, another rocket bears the Bloodhound name but without the warlike aura. This one aims to transport a human being, not an explosive charge, in an attempt upon both the land and low level aviation speed records at 1·4 times the sea level speed of sound. At the same time, it is generating data and methods for spin off into wider scientific and technological theatres and scientific computing is at its heart. [more]

11 October 2012

Fire, flood, pestilence and war

Statisticians, despite popular perceptions to the contrary, are as fond of play as the next person – and for a statistician there is no bigger playground than ecology. It starts off at the same size as the planet, layers dimensionally upwards downwards from human scale, has multiple expression in uncountably many scientific study domains, extends backwards and forwards in time. And throughout all of that, it is intrinsically statistical.
The surge of public environmental attention in the 1960s and 1970s after publication of Rachel Carson's Silent spring was a key factor in making me a statistician, and many others of my generation make the same admission. Hippy illusions gave way to irresistible glimpses into endless unexplored vistas of complexly related data ... and then came scientific computing to make the adventure feasible.
Ecology is, in conceptual essence, a statistical study of chreodic systems in perpetual flux with dynamic equilibrium frequently rearranged by catastrophic sheer planes.

  • Rachel Carson, Silent spring. 1962, Boston, Mass.: Houghton Mifflin.

17 August 2012

What are your thoughts?

“I used to think I was indecisive, but now I’m not so sure”. So runs one of the oldest entries in the bumper fun book of psychology jokes. It parallels one in the statistician’s equivalent volume: “statistics means never having to say you’re certain” (which, if you are too young to remember Ali McGraw and Ryan O’Neal[1] in Love story, plays off the strapline “Love means never having to say you’re sorry.”) And both have a serious echoes in professional practice: because psychology is an area in which data trees are always compromised by a forest of confounding factors, and samples can often be small. In most pure, classical psychology, to an even greater extent than in the physical sciences, only statistical analysis can (to mix metaphors) tease out the pure signal from the white noise with confidence.

Even in its pure and classic form, the concerns of even the most abstract psychological research are rarely far from pragmatic utility. Roland Bremond, a mathematical morphologist currently concerned with the very down to earth business of transport planning, neatly encapsulated the data analytic nature of the beast[2] in Saussurian terms: “La «réalité», du monde des objets, en psychologie expérimentale, est définie statistiquement comme l’ensemble des propriétés perçues partagées par tous les sujets «normaux» ... ... ... Autrement dit, la réalité est une propriété statistique, sur une population...” (In experimental psychology “reality”, the world of objects, is defined statistically as the set of properties shared by all perceived “normal” subjects ... ... ... or, to put it another way, reality is a statistical property of a population...)

Not that pure classical psychological research is the only strand or even the most evident. Without straying too far into monist and dualist controversies it’s fair to say that complete separation of mind from body was abandoned a long time ago, and psychology interpenetrates deeply with neuroscience and biochemistry amongst many other physical areas.

In its classic guise, especially in combination with modern computerised statistics, psychology is one of the areas of science where real, useful, original research can still be done by the solitary researcher without institutional funding or resources. In my freelance consultancy life I regularly work with research psychologists and frequently encounter such individuals, most of them right at home in one data analysis software package or another. In the run up to this article several lines of thought came from lone workers including the occasional pre-university student. [More]

15 June 2012

Navigating the sea of genes

Over my working life, statistical data analysis has explosively expanded in significance across every area of scientific endeavour. Over the last couple of decades, computerised methods have ceased to be an aid and become, instead, simply the way in which that statistical data analysis is done. Partly as a result, and partly as a driver, data set sizes in many fields have grown steadily. As with every successful application of technology, a ratchet effect has kicked in and data volumes in many fields have reached the point where manual methods are unthinkable.

Some areas of study, however, have seen their data mushroom more than others. Of those, few can match the expansion found in genetics, a field which has itself burgeoned alongside scientific computing over a similar time period. The drive to map whole genomes, in particular, generates data by the metaphorical ship load; an IT manager at one university quipped that “it’s called it the selfish gene for a reason, and not the one Dawkins gave: it selfishly consumes as much computational capacity as I can allocate, and then wants more”. [more]

14 April 2012

Painting by numbers

One of the staple exercises in statistics education is to make a light hearted foray into that old academic wrangle: were the Shakespeare plays and sonnets really written by William Shakespeare or by [insert your favourite candidate here]? Various metrics are analysed by whatever techniques are being taught, with a view to assessing similarity and difference. Real literary academics, of course, have visited the same methods with serious intent, and similar debates exist within the visual plastic arts. Was this unsigned painting produced by old master X, or by unknown Y?

A very well known example of such a dispute concerns the early 17th century Baroque painter Artemisia Gentileschi. In the last fifty years she has been rehabilitated in art history circles and is now widely recognised, but for centuries much of her work was wrongly attributed to others. The most frequent misattributions were to her father Orazio (which most experts now regard as blatantly ridiculous but which can be blamed on signatures) or to Michelangelo Merisi da Caravaggio who painted in a superficially similar style a generation earlier. Reattribution was in most cases visual by expert witnesses but in a few cases appeal was made to more objective means.

Disputes still occur, however, and as recently as January of this year a bitter disagreement between two national art institutions, with significant financial implications, was finally settled. The process is plastered with nondisclosure clauses, but people love to talk about their interesting cases... [more]


  • Illustration: Artemisia Gentileschi, La Pittura, c1638. Oil on canvas, 986mm × 752mm. Part of The Royal Collection.

13 April 2012

Perfecting pills and potions

“You are here to learn the subtle science and exact art of potion making ... the delicate power of liquids that creep through human veins, bewitching the mind, ensnaring the senses ... I can teach you how to bottle fame, brew glory, even stopper death.”[1]

J K Rowling's potions master, Professor Severus Snape, is talking of magic but his approach is in many ways more scientific (“there is little foolish wand waving here”, he warns) than muggle explanations for drug effects which were accepted professional consensus until quite recently in the scheme of things. In anything like the present day usage of the term, pharmacology is (give or take an argument or two over detail) a hundred and sixty five years old. For much of that short history, pharmacological data analysis was externally observational: the body was a black box, with the drug as input and detected effects as outputs. The current receptor based molecular approach, though it has its theoretical roots in the early twentieth century, only really dates from after the second world war. [more]


  1. Rowling, J.K., Harry Potter and the philosopher's stone. 1997, London: Bloomsbury. 0747532699 (cased) 0747532745 (pbk.).

19 February 2012

Making mince pies and MVD

Every year, in December, I embark on an epic quest to find mince pies. I have an addiction to these pastries, but also a particular dietary requirement which is not explicitly provided for by suppliers. I usually find a supply (and promptly buy a full gross before they can get away), but in the process I encounter the effects of what is, in the scheme of things, a trivial aspect of statistics in manufacturing to which I shall return below. When I last wrote[1] about this topic in Scientific Computing World, I was mostly concerned with the central rôle of data analysis in quality control. It plays a much wider and more diverse range of parts than that, however.

The fashion industry, for instance, is driven by statistical compromise. A dress maker uses over thirty metric descriptors, from height or waist circumference to the distance from neck to shoulder or armpit to hip. Most dresses, on the other hand, are bought on the basis of a single descriptor: a size, which in theory will always mean the same thing but in practice varies widely. In making this data reduction, a manufacturer needs to ensure that the best possible balance is struck between different customer shapes and perceptions. Most customers are going to find that a dress fits in one place, is too loose in another, and too tight in a third, but will only accept this within certain limits. Moving to a larger or smaller size will alleviate one problem while exacerbating another, and will also carry a psychological message about body shape. Getting the compromise wrong for a particular market will result in an exodus of customers to another manufacturer who has made shrewder decisions. These choices are heavily influenced by intuition and experience, but modelling on the basis of data analysis plays a large part, too. I know of one high street fashion supplier who regularly rents consulting time from a university department which maintains dedicated analytic software for the purpose. There are others which use in house statisticians running desktop software... [more]

26 December 2011

Difference of opinion

Today, as Ray Girvan notes at JSB, is Charles Babbage's 220th birthday. By a coincidence, I have just built a portal to a set of Babbage and Lovelace diary pages which Ray and I jointly wrote some years ago and with which I am still pleased ... a coincidence which seems like a flimsy enough excuse to shamelessly remind my readers of its existence.

21 December 2011

Cars, cows and carbon sinks

An externality, to an economist, is[1] "a side-effect or consequence ... which affects other parties without this being reflected in the cost of the goods or services involved". Externalities take all sorts of forms, and can be positive or negative, but over the past half century industrial pollution of the environment has become the primary exemplar.

More recently still, the focus has narrowed down to carbon based compounds whose costs are paid in a number of ways. The crudest direct health effects are usually localised, and become a matter for local legislation or lack of it; the atmospheric greenhouse effect is a global issue with no respect for human jurisdictional boundaries.

Attempts to deal with pollution almost always come down to mechanisms designed to convert an externality into a direct cost paid by the polluter, and carbon is no exception. A number of schemes exist to license carbon emission, with a market in which those who emit least sell permissions to those who emit most, thus exerting a direct proportional cost pressure on producers.

Whether this method is effective, and if so to what degree, is a subject of considerable political argument; but it remains the principle approach. Its use depends on quantification of emissions, which is neither simple nor straightforward. In practice, output is usually simplified from the full gamut of emitted substances (not all of them carbon based) to a single carbon dioxide equivalence figure, the product of mass and a radiative forcing factor which varies from substance to substance. But that still leaves an impractically large data acquisition and monitoring task.

The essence of statistical data analysis, always and everywhere, is generalisation from sample to population with a quantified level of confidence. Sometimes, as with extraterrestrial exploration in the last issue, this is because only tiny amounts of data can be captured and the maximum information must be squeezed from it. In the case of planetary emission levels the opposite is true: the available data volume is huge, and only a small fraction of it can be manageably handled. [more]


1. Oxford English Dictionary, Oxford University Press.

14 October 2011

Mars and the asteroids...

In the bizarrely nonsensical words from my schooldays, "Mary Voraciously Eats Mother's Jam Sandwiches Under No Protest". In case your own childhood did not include that particular mnemonic phrase, it represented the sequence of planets in order of distance outward from the sun.

Pluto has since been demoted, and new mnemonics have emerged, but that needn't trouble us here because the imaginative focus of interplanetary attention is now on Mother's Jam: that is, on Mars and Jupiter. Last year, US president Barack Obama envisaged a human landing on Mars in the mid 2030s and NASA's Ames Research Centre has jointly invested with DARPA in the idea of a one way Mars colonisation project. Russian plans over similar time frames include robotic exploration of Mars' moons. As you read this, NASA's Juno mission will be several weeks into its five year journey to Jupiter.

At a less romantic but perhaps more immediately practical level, there is also interest in the sweep of rocky space between them: the asteroid belt. On one level, it is a valuable scientific repository of "cosmological memory". At another, all exploration has, behind its heroic image, investment in the hope of economic return. The asteroids hold out the tantalising dreams of achieving that return well within a human lifetime; Mars within a century; Jupiter only in the much more distant future. Obama's vision for NASA includes not only the Mars mission but an asteroid ready heavy lift rocket design to be complete "no later than 2015", and the realities of returning from asteroid to earth orbit are trivial compared to Mars.

Mars has, of course, so far been subjected to more extensive examination than any other extraterrestrial target apart from Earth's own moon. A dozen or so programmes have, despite numerous failures, built up a knowledge base upon which projected US, European, Russian and Chinese successors plan to build over the next decade or so. The asteroids have mostly been studied remotely, usually in passing while on the way to somewhere else, but greater direct attention is now being paid to them. From an economic standpoint, they represent a potential resource for materials which would otherwise have to be lifted out of Earth's gravity well (and finite supply) at immense cost.

In all cases, however, before the economic return comes investment in study based upon huge programmes of data analysis. [More...]


Image: Orbital image of the Ma'adam Vallis flow channel, entering the Gusev crater at the top of the frame. [Source: NASA]

22 August 2011

Bringing technology to life

Biomimetic electromechanical prostheses are delivering the first generation of active replacement parts, but between biological inspiration and industrial delivery comes a lot of data analysis. [more]

09 June 2011

A healthy approach to data analysis

As this appears, by a happy piece of synchronicity from my point of view, the Wellcome Collection in the UK has on show an exhibition called Dirt: the filthy reality of everyday life. One exhibit in particular is of pivotal relevance to data analytic epidemiology: Dr John Snow’s so called "ghost map". In 1854, using what would today be described as data visualisation, Dr Snow plotted cases of cholera on a map of Soho, London. From the results he deduced that a water pump, was the source of infection. This was particularly impressive because water was not, at the time, suspected as a transmission vector and the pathogenic germ theory of disease had not become generally accepted. The local council decision to disable the pump was therefore, in the circumstances, a seminal act of faith in datacentric deduction over conventional wisdom.

Seemingly unlikely causation chains are often discovered by more sophisticated variations on Snow’s theme, emerging through statistical winnowing of gathered data. More than most data analytic areas, epidemiology can benefit from pooled work by numerous users at the sharp end of their practice as well as high level overviews, and data analysis is vital across that whole range. Those who have me in preparing this article include theatre nurses, general practice managers, country vets and hospital porters.

In a more recent high profile example, again involving cholera, an outbreak in Haiti after the devastating earthquake seems to have been traced¹ to a tragic “confluence of circumstances” arising from the aid effort itself. Identification of the apparent initial import vector didn’t require any sophisticated analysis in this case, but patterns of spatial spread within the country subsequent to that were a different matter. In an unfunded study of data from census and hospitalisation records (using Madonna software, widely used software from the University of California at Berkeley) Tuite and others² were able to model transmission in a way which “Despite limited surveillance data ... closely reproduces reported disease patterns”. [more]


  1. Cravioto, A., et al. Final Report of the Independent Panel of Experts on the Cholera Outbreak in Haiti. 2011, New York: United Nations News Service Section.
  2. Tuite, A.R., et al., Cholera Epidemic in Haiti, 2010: Using a Transmission Model to Explain Spatial Spread of Disease and Identify Optimal Control Interventions. Annals of internal medicine, 2011. 154(8).

I would like to thank Dr Brian Corden for invaluable help in assessing an item which, as a result of his advice, was not eventually used. He thus saved me from making a fool of myself through lack of confidence in my own judgment