Showing posts with label Data. Show all posts
Showing posts with label Data. Show all posts

08 November 2013

Too big for its boots

In a delightful way, from a writing perspective, my last few data analysis topics have first synergised with one another and then led naturally on to the consideration of Big Data.
What exactly is "big data"? The answer, it should come as no particular surprise to hear, is "it depends". As a broad, rough and ready definition, it means data in sufficient volume, complexity and velocity to present practical problems in storage, management, curation and analysis within a reasonable time scale. In other words, data which becomes, or at least threatens to become in a specific context, too dense, too rapidly acquired and too various to handle. Clearly, specific contexts will vary from case to case and over time (technology continuously upgrades our ability to manage data as well as generating it in greater volume) but broadly speaking the gap remains – and seems likely to remain in the immediate future. The poster boys and girls of big data, in this respect, are the likes of genomics, social research, astronomy and the Large Hadron Collider (LHC) whose unmanaged gross sensor output would be around fifty zettabytes per day.
There are other thorny issues besides the technicalities of computing. Some of them concern research ethics: to what extent, for example, is it justifiable to use big data gathered for other purposes (for example, from health, telecommunications, credit card usage or social networking) in ways to which the subjects did not give consent? Janet Currie (to mention only one recent example amongst many) suggests a stark tightrope with her "Big data vs big brother" consideration of large scale pædiatric studies. Others are more of concern to statisticians like me: there is a tendency for the sheer density of data available to obscure the idea of a representative sample– and a billion unbalanced data points can actually give much less reliable results than thirty well selected ones.
Conversely, however, big data can also be defined in terms not of problems but of opportunity. Big data approaches open up the opportunity to explore very small but crucial effects. They can be used to validate (or otherwise) smaller and more focussed data collection, as for instance in Ansolabehere and Hersh’s study [1] of survey misreporting. As technology gives us expanding data capture capabilities at ever finer levels of resolution, all areas of scientific endeavour are becoming increasingly data intensive. That means (in principle, at least) knowing the nature of our studies in greater detail than statisticians of my generation could ever have dreamed. A while back, to look at the smaller end of the scale, I mentioned [2] the example of an automated entomological field study régime simultaneously sampling two thousand variables at a resolution of several hundred cases per second. That’s not, by any stretch of the imagination, in LHC territory but it’s big enough data to make significant call on a one terabyte portable hard drive. It’s also a goldmine opportunity for small team or even individual study of phenomena which would not long ago have been beyond the reach of even the largest government funded programme: big data has revolutionised small science.
There is, in any case, no going back; big data is here to stay – and to grow ever bigger, because it can. Like all progress, it’s a double edged sword and the trick as always is to manage the obstacles in ways which deliver the prize. [more]

[1] Ansolabehere, S. and E. Hersh, Validation: "What Big Data Reveal About Survey Misreporting and the Real Electorate". Political Analysis, 2012. 20(4): p. 437-459.

[2] Grant, F., "Retrieving data day queries", in Scientific Computing World. 2013, Europa Science: Cambridge. p. 10-12..

21 February 2013

Omnilingual word spotting

In yesterday's out take from my recent “Joy of text” piece in Scientific Computing World, I mentioned in passing the term “word spotting”.
Word spotting looks for visually similar discrete components within a text, and classifies them using statistical comparisons. In approximate human terms, it is treating words (or combinations of words) as ideograms rather than as phonetic constructs.
Word spotting is not limited to written or printed material, though that is the context with which I'm concerned here: it also applies, for example, to speech recognition. Nor, intriguingly, is it necessarily limited to words of known meaning; it can equally well be applied to semantic units of entirely unknown signification. It could, as an extreme example, be used to analyse the manuscripts from a lost extraterrestrial civilisation in H Beam Piper's classic science fiction story Omnilingual.
Reverse the Omnilingual example to imagine a hypothetical extraterrestrial archaeologist trying to study post apocalyptic remains of our own cultures. It will be obvious from context that certain signification units are associated with the physical sciences: "volts" on electrical signs and appliances, just to pick one.
Our xenoarchaeologist (who may not have alphabetic scription systems, or even vocalisation, and certainly cannot assume that each letter represents a sound) looks at a mass of textual matter whose content, subject matter, purpose and reliability are unknowable. Where should attention be concentrated? The only certain knowledge is that words are identifiable visual entities found in isolation and occurring with spatial separation in books.
Word spotting, with no semantic assumptions, quickly shows that some books contain many instances of the visually related signifiers "volts" "volt", "voltage", "voltmeter" and so on, suggesting that those sections contain material related to electricity; others do not. There will, of course, be false hits such as "Voltaire" and "revolt", but as one discriminator amongst others in a multiple sieving process it would nevertheless be invaluable.
Handwritten notes and journals would be less amenable than printed books (in my own handwriting, for instance, computer transcription systems have trouble separating "volt" from "bolt" and sometimes even "void"), but could still be sieved using multiple discriminators in the same way.

  • Piper, H.B., Omnilingual, in Astounding Science Fiction. 1957, Dell Magazines: Northwalk CT.

20 February 2013

Making sense of the census

In preparing an article on any topic, there is often more good stuff than makes it into the final edit.
In the case of my recent “Joy of text” piece, about text analysis for Scientific Computing World, one of the out takes was a prototype content-based image retrieval system framework developed for census searches by Kenton McHenry and his ISDA group at NCSA. The problem to be solved is computerised searching of large volume handwritten census returns.
A user inputs a handwritten query – I might, for instance, write "Grant". The system derives a numerical feature vector which describes that input, then seeks occurrences of similar vectors within the image database.
The system is designed to self-validate, by recording which returned entities are selected by the user. I will not, for example, select false hits such as "Grand" or "Ghent" and the system will note which results I do or do not follow up. Over time, other Grants will make similar decisions and increase the system's confidence in selecting some hits for return and not others; gradually, those writing “Grant” as their query will see fewer and fewer offers of documents containing similar looking words.
The computer analytic process behind all this is progressive.
First, the lines and boxes on the census forms are used to carve up the content into image segments (for instance, surname will be in a box at the same location on each form and will become a data entity). Each segment is then converted into a numerical feature vector representing its appearance, and similar feature vectors are grouped hierarchically. Two million XSEDE (Extreme Science and Engineering Discovery Environment) CPU hours have been requested for initial record processing.
When the search query is entered, word spotting is used to compare its vector with those stored in the database, seeking matches within statistically defined limits of similarity. The search is not a blind one from the beginning of the database through seventy billion image segments to the end; the hierarchical grouping guides greatly reduces the number of entities which need to be compared.

15 February 2013

The joy of text

It's a mundane truism, not normally worth mentioning, that words and phrases as signification units in natural language have only the fuzziest of relations to that which they signify. It is, nevertheless, a live issue for the many researchers attempting to computerise data analytic activity using text as raw material. It's also a truism of which I have been reminded afresh as I discussed the topic with practitioners and consumers of textual analysis, no two of whom used the term in exactly the way.
Strictly speaking, textual analysis describes a social sciences methodology for examining and categorising communication content. In practice, though, it is widely used to cover a range of activities in which unstructured or partially structured textual material is submitted to rigorous analytic treatment. What they all have in common is a desire to wrestle the petabytes of potentially valuable information locked up in an ever inflating text reservoir (blogs, books, chat rooms, clinical notes, departmental minutes, emails, field journals, lab notebooks, patents, reports, specification sheets, web sites and a million other sources) into a form which is susceptible to useful, objective data analytic treatment. Temis, of whom more later, have on their website a headline which sums it up neatly: "Big data issue #1: a lot of content and no insights". Text mining, the consequent knowledge bases, and analysis of the results have become a major component of biomedical and pharmaceutical research.
For our purposes here, I have taken it to mean analysis whose purpose is to extract scientific value from texts, to examine those texts scientifically, or some combination of the two. [More]

20 January 2013

Ode to OCR

Preparing an article on text analysis for the upcoming issue of Scientific Computing World (watch this space) I've been revisiting the various ways of getting physical text into digital form.
Optical Character Recognition (OCR) is the workhorse of text transcription and, while we all grumble about its very real real short comings shortcomings, it really does a remarkably good job, most of the time, of rendering graphic images of printed fonts into digitised text for analysis.
Even at the lowliest manual level, and with all its admitted faults, OCR is a useful tool. A colleague and I recently had to add a six hundred and fifty page eighteenth century text to a digitised textbase for analysis. The rare and valuable paper original which we located was in a library, and could not be removed. Filing a request for the digitisation to be carried out would take weeks. With the consent of the library we used a smartphone, a ten year old copy of ABBYY FineReader 5 (now in release 11, and correspondingly more developed, as part of a software range for different text tasks) and a netbook. Even allowing for manual error correction, we had our validated data within three hours and the library added a copy to its own digital records. A similarly sized text already available as graphic only PDF was transferred more quickly still.

04 July 2005

Survey of 260 British teenagers

Data set from self declaration questionnaires, 2005.

This data set has no been shifted into an Excel spreadsheet file which can be downloaded by clicking here.

08 May 2001

Aegean clays data set

This dataset is part of a collection assembled during research by Dr Ioannis Liritzis, Professor of Archaeometry at the University of the Aegean at Rhodes, and is used by his kind permission.
Spectroscopic analysis of the clays yielded figures for varying levels of different elements (this small subset concentrates mainly on metals) in the composition of each clay.

The data were gathered as part of an investigation into historical trading links between Mediterranean communities, in which the composition of clays used in manufacture of dated pot shards from separate sites were analysed for similarity/difference.
The data are posted here in four tables.
I recommend the Mediterranean Archæology and Archæometry journal  (started by Prof. Liritzis and still under his overall editorship).

Aegean clays data set (Table 1)

Data set: Relative element abundance in Aegean clay samples
Courtesy of Dr I Liritsis, Dept of Archaeometry, University of the Aegean at Rhodes

Site Grp Ti Fe Mn Ca K Cr Ni Cu
SAR 5 S 2031 22061 668 121768 14864 286 55 334
YCEM3 Y 3444 26897 789 20257 16310 120 1625
YALINK Y 4044 36113 161 16638 12598 142 129
YALD2 Y 3233 33398 1127 12558 13938 175 142 122
YALD5 Y 3755 34035 15895 14714 291 57
PERG1 P 4266 38805 572 28642 10768 1334
SAR 282 S 1522 16487 539 125969 18870 429 83 88
YAL7CP Y 3355 31779 11083 13102 203 79
YD2 Y 4577 45562 700 26151 13546 256
YALD6 Y 3277 37732 668 16591 14284 205 165 217
YB1S Y 3711 31622 467 9240 19744 145 96 124
SAR 240 S 1463 15448 346 158804 18546 289 94
YALD1 Y 3689 44367 773 18791 17756 144 181
SAR 290 S 1389 15281 403 122050 33968 296 81
YAL3NK Y 2454 35197 451 13188 14141 468 176 125
YZ2 Y 3933 43295 757 19947 13087 136 112 415
SAR 404 S 589 7071 628 169858 23599 179
SAR 180 S 1327 13248 725 138659 21566 245 138
SAR 210 S 1085 11796 386 160279 12756 298 75
YCEM1 Y 3866 40335 403 27316 14699 529
YCEM2 Y 3766 39296 789 19439 11679 119 735
SAR 120 S 1749 23256 338 71562 15083 324 156 146
YZ3 Y 2900 39196 475 61062 14940 208 249
SAR 190 S 1211 12533 531 179230 14036 172 64
YAL5NK Y 4555 41363 692 21225 10429 79 204
YAL8CP Y 3844 39899 314 19383 17733 200 103
YZ1 Y 3200 34069 346 13592 16205 338
SAR 410 S 852 8255 652 153540 22379 182 65
SAR 105 S 2308 23937 531 112151 17327 286 123 189
SAR 93 S 1549 17023 700 142015 15007 388 104 217

Aegean clays data set (Table 2)

Data set: Relative element abundance in Aegean clay samples
Courtesy of Dr I Liritsis, Dept of Archaeometry, University of the Aegean at Rhodes
 
 Site Grp Zn As Sr Zr Mo Pb Rb Ba
YB1 Y 620 71 458 191 3 33 85 1009
YALDNK Y 149 31 449 230 45 667
SAR 253 S 234 195 82 7 43 96
SAR 417 S 214 209 64 9 720 57 67
YAL6CP Y 79 49 728 173 70 511
SAR 392 S 347 338 56 7 216 11 104
SAR 80 S 615 86 196 76 8 66 112
YB2 Y 149 72 206 184 136 442
SAR 20 S 319 29 213 91 5 37 119
SAR 152 S 251 52 254 78 6 15 55 103
YAL4NK Y 144 102 1074 173 101 1129
SAR 345 S 244 238 67 5 36 115
YD1 Y 157 52 550 219 73 474
SAR 308 S 233 42 229 63 4 57 107
SAR 303 S 255 221 63 8 57 109
SAR 316 S 251 55 176 92 8 21 67 110
SAR 357 S 287 302 99 2 27 45 136
SAR 324 S 324 42 181 95 5 46 114
YNZ2 Y 324 399 129 75 1027
SAR 265 S 245 66 239 85 3 49 133
PERG2 P 506 97 700 225 4 81 1041
SAR 377 S 293 72 196 95 5 31 46 105
YAL2NK Y 115 60 476 200 5 82 634
YALD3 Y 144 52 544 176 88 562
SAR 48 S 544 71 257 108 6 53 175
YAL9TG Y 116 45 725 210 3 36 477
SAR 366 S 224 63 302 76 8 60 105
SAR 268 S 222 55 220 82 5 54 101
YNB1 Y 663 65 529 209 19 61 882
YA1 Y 285 40 582 166 4 15 71 990
YALD4 Y 87 86 426 209 74 806
SAR 60 S 461 85 163 113 5 32 66 127
SAR 135 S 298 120 198 83 5 43 110
SAR 232 S 179 43 308 59 6 40 101
SAR 386 S 404 274 75 4 197 36 191

Aegean clays data set (Table 3)

Data set: Relative element abundance in Aegean clay samples
Courtesy of Dr I Liritsis, Dept of Archaeometry, University of the Aegean at Rhodes
 
 Site Grp Ti Fe Mn Ca K Cr Ni Cu
YB1 Y 4244 42278 845 19674 20723 110 1118
YALDNK Y 5133 41552 547 19975 11815 105 53 179
SAR 253 S 1318 16576 113383 22394 268 33
SAR 417 S 1102 12343 322 106229 25933 355 84
YAL6CP Y 3300 32683 652 16666 14405 211 231
SAR 392 S 1183 8847 668 208492 14382 350 115
SAR 80 S 1620 19212 443 124353 17643 295 171 163
YB2 Y 3244 25881 555 11853 23893 250 164
SAR 20 S 1657 21189 475 82579 13373 181 226
SAR 152 S 1502 16319 459 113054 22462 309 82 151
YAL4NK Y 3988 45216 547 19646 18501 73 151
SAR 345 S 1132 12466 298 172528 16574 238 105
YD1 Y 4322 40469 483 23077 13351 122 308
SAR 308 S 1354 15582 451 117500 31061 205 96 125
SAR 303 S 1267 13270 386 115300 39758 343 99
SAR 316 S 1479 17861 684 96030 34661 323 70 85
SAR 357 S 1780 20854 998 163250 16318 299 44
SAR 324 S 1920 21268 950 95692 17436 396 92
YNZ2 Y 3077 33030 708 14767 15677 297 103 472
SAR 265 S 1664 16632 668 135031 27899 368 78
PERG2 P 5022 40703 966 19120 19826 722
SAR 377 S 1597 17738 362 116250 21912 351 97 99
YAL2NK Y 2764 29210 17860 13065 244 125
YALD3 Y 4055 39073 596 16582 17673 125 167
SAR 48 S 2511 25993 90983 23448 390 180 306
YAL9TG Y 3589 36191 692 24609 9766 187
SAR 366 S 1960 18185 1063 140286 20926 254 108 96
SAR 268 S 1233 15459 386 116353 26332 262 57
YNB1 Y 4533 43306 797 22795 11852 1177
YA1 Y 4200 36425 765 16723 18087 224 102 464
YALD4 Y 3011 33499 282 15134 14571 111 101
SAR 60 S 1896 25311 161 59718 21596 317 230
SAR 135 S 1263 15750 370 114661 17522 203 85
SAR 232 S 1284 14968 451 180781 20226 284 163
SAR 386 S 1387 14454 700 176711 18185 200 66