Monday, 7 March 2011

Visualising (dirty) data from data.gov.uk using the Dataset Publishing Language (DSPL)

A fortnight ago the Dataset Publishing Language (DSPL) was launched by the Public Data Team at Google. DSPL is an XML-based language to support the generation of rich and interactive data visualisations using the Public Data Explorer, Google's hitherto closed visualisation tool. The XML is used to describe the dataset, including informational metadata like descriptions of measures and metrics, as well as structural metadata such as relations between tables. The completed DSPL XML is then uploaded to the Public Data Explorer in a 'dataset bundle' containing a set of CSV files containing the data of the dataset.
I decided to take the DSPL for a spin using data gleaned from data.gov.uk and visualised data pertaining to UK higher education income and expenditure in the years up to 2008 and 2009. This process was a little fidgety, primarily for reasons to be discussed in a moment; but it was also fidgety owing to the demands of the DSPL and the seemingly temperamental nature of the Public Data Explorer. (These technical issues are something the Public Data Team is resolving). The dataset can be visited and enjoyed as a bar graph, bubble chart or line graph, with dimensions selected from the left-hand column and temporal dimensions under the X axis. Bubble metrics in the bubble graph can be toggled in the top right-hand corner. Note that all values are shown in units of 1000 GBP, and where necessary rounded to the nearest 1000 GBP. Screenshots are above and below.

These data visualisations look very good indeed, and this will no doubt be a useful resource for many. But I can't help wondering if it's all too much pain for too little gain. The dataset I used is relatively simple but it still required 140 lines of XML and an endless amount of tinkering with the original data. So unless you have a large, pristine dataset which is to form the focus of a keynote presentation at an important conference (such as Prof. Hans Rosling), it is difficult to see whether it is worth the effort. Added to which, ironing out errors in the DSPL is arduous because the Public Data Explorer is only clever enough to tell you that there is an error, not where the error might be. This is all very frustrating when your XML is well-formed, validates, and your CSV files appear kosher. Again, the Public Data Team is working hard so things should improve soon. Which brings me back to the principal reason why the whole process was fidgety: data.gov.uk.

Data.gov.uk was launched a year ago by Tim Berners-Lee on behalf of the UK government. You can read about the background in your own time. Suffice to say, the raison d'etre of data.gov.uk is to publish government datasets in an open, structured and interoperable way thus stimulating new and "economically and socially valuable applications". As it currently stands, data.gov.uk does not come close to achieving this. It is not until you delve beneath the surface (as I did for the dataset above) that you appreciate what data.gov.uk actually provides is almost the opposite: closed, unstructured and un-interoperable data! A resource like this should be based – in an ideal world – on RDF or XML, with CSV the preferred option for those unfamiliar, unwilling or unable to provide something better. But it should not be a repository for virtually every file format known to human-kind, with contents structured in an arbitrary manner.

Identifying a suitable dataset for my DSPL experiments was exhausting. PDF files are commonplace; some "datasets" are simply empty or broken, or are simply bits of information (e.g. reports). Even if you are lucky enough to find a CSV compliant dataset (and don't expect any RDF or XML), it will inevitably be dirty and require significant time to render it usable, hence why my experiences were fidgety. All of these frustrations appear to be shared by developers that post on the data.gov.uk forum. To be sure the data is "open" insofar as UK citizens can visit data.gov.uk, view data and hold public officials to account. However, it's the data.gov.uk logo (three linked orbs) - which is almost identical to the old Semantic Web logo - that seduces one into thinking data.gov.uk it is a rich source of structured, interoperable, open data. None of this is entirely fair because data.gov.uk does have a page on Linked Data, and it does provide some useful RDF on MPs, legislation, etc. and some SPARQL endpoints; but in the grand scheme of 'all-things-data.gov.uk' it constitutes a very small proportion of what data.gov.uk actually provides. And all of this is very depressing. It increases barriers, alienates the developers and data enthusiasts, and will ultimately fail to reach the objective: "economically and socially valuable applications".

Friday, 4 February 2011

From the bottom up: growing the Semantic Web


RDFa is essentially a form of XHTML which incorporates a variety of RDF attributes and can adhere to the RDF data model. It's generally far less detailed than standalone RDF; but it does the trick for most web pages. (See earlier postings for some further information). As a Semantic Web aficionado, I was therefore pleased to read Peter Mika's recent blog concerning the recent growth of RDFa deployment.

Mika is a researcher at Yahoo! Research specialising in semantic search. His research is widely published and his role at Yahoo! has enabled him to analyse the growth of RDFa on the surface web. Mika's analysis charts the evolution of certain microformats and RDFa on the web, as percentage of all web pages, as indexed by Yahoo! Search. This includes over 12 billion web pages. The data collection was conducted at at three different time points over the past two years, thus allowing growth to be charted. There are some caveats with the data, but overall Mika found the following:
"The data shows that the usage of RDFa has increased 510% between March, 2009 and October, 2010, from 0.6% of webpages to 3.6% of webpages (or 430 million webpages in our sample of 12 billion). This is largely thanks to the efforts of the folks at Yahoo! (SearchMonkey), Google (Rich Snippets) and Facebook (Open Graph), all of whom recommend the usage of RDFa. The deployment of microformats has not advanced significantly in the same period, except for the hAtom microformat."
510% growth in 18 months??? What an incredible statistic. It doesn't matter that academics, the BBC, The Guardian, NY Times, data.gov.uk, DBpedia, digital libraries, the CIA, etc. participate in the Semantic Web using 'pure' RDF; it takes Yahoo! (which was leading the search engines in the use of RDFa) and Google (which recently acquired Metaweb Technologies) to motivate ordinary web developers to use RDFa. All of this is very motivating. I wonder where we'll be by October 2011? Hopefully Mika will update his blog sometime in the autumn!

Monday, 24 January 2011

Delicious: an obituary of sorts

University bureaucracy was swallowing up my time when this news broke, otherwise I would have commented sooner... But the leaked announcement by Yahoo! that it has decided to 'sunset' Delicious was momentous, and although Delicious may have a future elsewhere, the news remains significant. It's significant because – along with Flickr, which Yahoo! also owns – Delicious was one of the first services which epitomised the new and mysterious 'Web 2.0' concept, when it emerged in 2004.

In 2004 the concept of organising and sharing links on the Web was fresh and new, and Delicious was really the first to offer an innovative solution to save, organise and share bookmarks with friends. Delicious popularised the use of bookmarklets and practically coined the term 'social bookmarking'. It was probably the first Web 2.0 service to make a domain hack cool and not cheap looking (it was http://del.icio.us/ until 2008, and actually called itself del.icio.us initially).

More interestingly, Delicious popularised tagging and – probably more than any other service at that time – launched an avenue of functionality known as 'social tagging'. In the deluge of social tagging research papers that have been published since Delicious was launched, few will not have cited Delicious within its introduction. And, of course, social tagging - or social bookmarking, or collaborative tagging - has come to be one of the defining aspects of Web 2.0. Social tagging has sent shock waves throughout the Web, influencing the design of subsequent social media services and discovery tools, such as digital libraries. Even the ubiquitous (and infamous) tag cloud - which sister service Flickr invented - was adopted by Delicious and rendered infinitely more useful with the uniquely identifiable resources it and its users curated.

And all of the aforementioned was why Yahoo! decided to acquire Delicious in late 2005 for circa $30 million, in what commentators noted as the first attempt by a Web 1.0 company to jump on to the Web 2.0 bandwagon. Of course, since 2005 many of the original Web 2.0 names have found a Web 1.0 home in which to evolve (e.g. Delicious and Flickr @ Yahoo!; YouTube, Picasa, Blogger @ Google; MySpace @ News Corp; etc.). To be sure, Yahoo! paid too much for Delicious; but Yahoo! weren't buying the technology (which at the time of purchase wasn't particularly complex). They were buying the brand, its users and the Web 2.0 kudos. However, Yahoo! failed to capitalise on all of that, and even failed to harness the underlying technology to make Delicious a household name. One would have thought that an injection of Yahoo! R&D would have made Delicious the most innovative social bookmaking service available, with plenty of horizontal integration with other Yahoo! products. Far from innovating, Delicious has been static since its acquisition. Five years on there are numerous social bookmarking services, virtually all of which are more innovative, more exciting and ultimately more useful.

It is an end of an era to be sure. No-one talks about 'Web 2.0' any more because there is no Web 2.0 to point to, and the demise of Delicious is an example of this. Web 2.0 is now about a handful of social media behemoths. To some extent the 'sunsetting' of Delicious is a comment on the utility of tagging and the value that can be mined from tags, and, ergo, the money that can be made from them. Tightening the use of tags is something which has attracted more attention recently from the CommonTag initiative and rival social bookmarking services such as Faviki and ZigTag. TechCrunch suggest that Yahoo! could have made money from Delicious if they had wanted to and that organisational issues prevented Delicious from being profitable. Perhaps they are right. Only a small team would have been required - but it can't have been easy to make money otherwise Yahoo! would have done it. The moment for Delicious is now gone... And this is an obituary of sorts: can you see any company thinking Delicious is a good investment?

Monday, 15 November 2010

New undergraduate degree programme: BSc Business Communications at LJMU

This blog tends to focus on research and often comments on how technological developments will alter the management of information and the computation of data. Occasionally, however, we also discuss issues within the undergraduate and postgraduate degrees (and modules) our team happens to deliver. It is therefore worthwhile announcing to all those who read the blog that our team has launched a new undergraduate degree programme for 2011: BSc Business Communications.

In BSc Business Communications at LJMU (UCAS Code: N102) students will study the strategic importance of communication, information and technology, and the role these play in the modern business organisation. Further information on the new programme can be found at our standalone BSc Business Communications website, the official LJMU BSc Business Communications website, or our Facebook group (BSc Business Communications at LJMU). BSc Business Communications is recruiting now for 2011/2012.

Tuesday, 2 November 2010

Crowd-sourcing faceted information retrieval

This blog has witnessed the demise of several search engines, all of which have attempted to challenge the supremacy of the big innovators - and I would tend to include Yahoo! and Bing before the obvious market leader. Yesterday it was the turn of Blekko to be the next Cuil. Or is it?

Blekko presents a fresh attempt to move web search forward, using a style of retrieval which has hitherto only been successful in systems based on pre-coordinated indexes and combining it with crowd-sourcing techniques. Interestingly, Rich Skrenta - co-founder of Blekko - was also a principal founder of the Dmoz project. Remember Dmoz? When I worked on BUBL years and years ago, I recall considering Dmoz to be an inferior beast. But it remains alive and kicking – and remains popular and relevant to modern web developments with weekly RDF dumps made of its rich, categorised, crowd-sourced content for Linked Data purposes. BUBL, on the other hand, has been static for years.

Flirting with taxonomical organisation and categorisation with Dmoz (as well as crowd-sourcing) has obviously influenced the Blekko approach to search. Blekko provides innovation in retrieval by enabling users to define their very own vertical search indexes using so-called 'slashtags', thus (essentially) providing a quasi form of faceted search. The advantage of this approach is that using a particular slashtag (or facet, if you prefer) in a query increases precision by removing 'irrelevant' results associated with different meanings of the search query terms. Sounds good, eh? Ranganathan would be salivating at such functionality in automatic indexing! To provide some form of critical mass, Blekko has provided hundreds of slashtags that can be used straight away; but the future of slashtags depends on users creating their own, which will be screened by Blekko before being added to their publicly available slashtags list. Blekko users can also assist in weeding out poor results and any erroneous slashtags results (see the video below) thus contributing to the improved precision Blekko purports to have and maintaining slashtag efficacy. In fact, Skrenta proposes that the Blekko approach will improve precision in the longer term. Says Skrenta on the BBC dot.Maggie blog:
"The only way to fix this [precision problem] is to bring back large-scale human curation to search combined with strong algorithms. You have to put people into the mix […] Crowdsourcing is the only way we will be able to allow search to scale to the ever-growing web".
Let's look at a typical Blekko query. I am interested in the new Microsoft Windows mobile OS, and in bona fide reviews of the new OS. Moreover, since I am tech savvy and will have read many reviews, I am only interested in reviews published recently (i.e. within the past two weeks, or so). In Blekko we can search like so…

"windows mobile 7" /tech-reviews /date

…where the /tech-reviews slashtag limits results to genuine reviews published in the technology press and/or associated websites, and the /date slashtag orders the results by date. It works, and works spectacularly well. Skrenta sticks two fingers up at his competitors when in the Blekko promotional video he quips, "Try doing this [type of] search anywhere else!" Blekko provides 'Five use cases where slashtags shine' which - although only using one slashtag - illustrate how the approach can be used in a variety of different queries. Of course, Blekko can still be used like a conventional search engine, e.g. enter a query and get results ranked according to the Blekko algorithm. And on this count – using my own personal 'search engine test queries' - Blekko appears to rank relevant results sensibly and index pages which other search engines either ignore or, if they do index them, normally drown in spam (spam results which these engines rank as more relevant).

There is a lot to admire about Blekko. Aside from an innovative approach to information retrieval, there is also a commitment to algorithm openness and transparency which SEO people will be pleased about; but I worry that while a Blekko slashtag search is innovative and useful, most users will approach Blekko as another search engine rather than buying into the importance of slashtags and, in doing so will not hang around long enough to 'get it' (even though I intend to...). Indeed, to some extent Blekko has more in common with command line searching of the online databases in the days of yore. There are also some teething troubles which rigorous testing can reveal. But there are reasons to be hopeful. Blekko is presumably hoping to promote slashtag popularity and have users following slashtags just as users follow Twitter groups, thus driving website traffic and presumably advertising. Being the owner of that slashtag could be useful, but also highly profitable, even if Blekko remains small.


blekko: how to slash the web from blekko on Vimeo.

Wednesday, 22 September 2010

New Google

My interest in information retrieval means that subscribing to search engine blogs (among other things) is essential. The most active blog to which I subscribe is the Official Google Blog. According to Google, the OGB provides "insights from Googlers into our products, technology, and the Google culture". More simply, the OGB is the place to look for developments in search, particularly those which Google wants to shout about.

There was a time (probably around two years ago) when updates to the OGB occurred every other week, and often the receipt of the RSS feed would compel me to post to this blog, such were the gravity of OGB announcements (see this, this and this, for example). However, in the past six months the OGB has been in overdrive. Almost every day a huge Google announcement is made on the OGB, whether it's the announcement of Google Instant or significant developments to Google Docs. Enter Google New, a new dedicated website to find all things new from Google. Here's the rationale from Google as published – yup, you guessed it – on the OGB:

"If it seems to you like every day Google releases a new product or feature, well, it seems like that to us too. The central place we tell you about most of these is through the official Google Blog Network [...] But if you want to keep up just with what’s new (or even just what Google does besides search), you’ll want to know about Google New. A few of us had a 20 percent project idea: create a single destination called Google New where people could find the latest product and feature launches from Google. It’s designed to pull in just those posts from various blogs."

Makes sense I suppose, eh?

Thursday, 9 September 2010

Web Teaching Day - 6 Sep 2010

On Monday, 6 Sept 2010 I attended a Web Teaching Day organised by Richard Eskins from Manchester Metropolitan University (his blog). We do a fair amount of web teaching in this group so I thought it would be useful to go along. Web teaching is undertaken by the Computing Department or the Art / Design department in most Universities and our courses tend to be very business orientated.

In many ways the conversations I had reminded me of those we have regarding Information Systems at John Moores. There's a problem relating to the range of skills required from basic technical skills, through design skills to high level inter personal skills. Our aim is to produce a "hybrid" graduate combining business with systems / technical skills. There's huge demand in industry for these graduates and our best students command very high salaries but students find the work hard and it difficult to recruit good students.

Web design / development courses have very similar aims.

Some of the highlights of the day:

Chris Mills from Opera talked about the Opera Web Standards Curriculum. He's been producing teaching material for students all of which is freely available on the internet. He also talked about Mozilla's P2PU (Peer to Peer University) programme - School of Webcraft which is aimed at delivering and assessing these skills. He's co-author on Interact with Web Standards: A Holistic Approach to Web Design (Voices That Matter).

David Watson from Greenwich University talked about the course he designed and runs - MA in Web Design and Content Planning. They've produced their own site to support the course. One interesting point he made is that he reckoned that this site had increased applicants to the course significantly, they now have 60 applicants for 20 places whereas before they struggled for numbers. He published his presentation here. He had some interesting observations on setting up and running the course. In particular don't depend on the University to market and recruit students to your course, you may end up with no students.

Christopher Murphy and Nik Persson (also known as the Web Standardistas) talked about their course BSc (Hons) Interactive Multimedia Design at Ulster University and the issues involved in delivering to undergraduates. I particularly liked their use of the nerd (Bill Gates) - designer (Steve Jobs) continuum to describe the difficulties of being a web builder and how you need so many disparate skills along this path. Here are some of the tools they recommend. Their book is HTML and CSS Web Standards Solutions: A Web Standardistas' Approach.

Aesha Zafar from the BBC talked about the new developments in Manchester and, in particular, the jobs that will be created there and Nicola Critchlow talked about the gap between industry's needs and graduates being produced by Universities (which is large and getting larger, nothing new there).

Finally, Andy Clarke, a freelance designer led a group discussion and chat at the end.

So what skills does a graduate from a web design / development course need? This is my list based on the nerd - designer continuum:

  • databases
  • server side programming languages (PHP seems to be in vogue though there are others)
  • Javascript
  • CSS
  • HTML including web development tools
  • graphics
  • design
  • people skills

It was an excellent day with lots of really inspiring speakers and it really got me fired up about the possibilities of delivering a web design / development course at John Moores. I don't believe that the course I want to offer exists here (though that's based on absolutely no research whatsoever!).

Our team has really strong skills in databases, programming, HTML, CSS and Javascript and we teach most of these skills at various levels. The people skills elements are taught throughout all our courses and is embedded in all JMU programmes via the World of Work (WoW) programme.

Our weakness is in design / graphics, however, Liverpool School of Art & Design has huge experience in areas such as graphic design and digital media.

So, here's a great opportunity to collaborate on a new course in an area that is growing in popularity.

Monday, 30 August 2010

Musical experiments with HTML5

The Official Google blog has just announced an HTML5 Chrome Experiment in association with Canadian indie rock band, Arcade Fire. This experiment appears to function as a marketing exercise for both Chrome and Arcade Fire; although it does also demonstrate that Google has a commitment to HTML5 (and it appears to be part of a wider partnership with Arcade Fire, as the video below indicates).

HTML5 is still currently under development but is the next major revision of the HTML standard (as distinct from the recent incorporation of RDF, i.e. XHTML+RDFa). HTML5 will still be optimised for structuring and presenting content on the Web; however, it includes numerous new elements to better incorporate multimedia (which is currently heavily dependent on third party plug-ins), drag and drop functionality, improved support for semantic microdata, among many, many other things...

The Chrome Experiment entitled, 'The Wilderness Downtown', uses a variety of HTML5 building blocks. In their words:
"Choreographed windows, interactive flocking, custom rendered maps, real-time compositing, procedural drawing, 3D canvas rendering... this Chrome Experiment has them all. "The Wilderness Downtown" is an interactive interpretation of Arcade Fire's song "We Used To Wait" and was built entirely with the latest open web technologies, including HTML5 video, audio, and canvas."
Being an 'experiment' it can be a little over the top, and I suppose it isn't an accurate reflection of how HTML5 will be used in practice. Nevertheless, it is certainly worth checking out - and I was quite impressed with canvas. An HTML5 compliant browser is required, as well as some time (it took 7 minutes to load!!!).

Monday, 23 August 2010

Jimmy Reid and the public library: an education like no other

Jimmy Reid was laid to rest last week. The obituaries have been plentiful and praising. As someone who is interested in the industrial history of Britain, I have always been especially interested in the industrial heritage of my home town of Glasgow (as well as my adopted home of Liverpool), and my special interest in Glasgow shipbuilding made Jimmy Reid's passing all the more sad...

Poster of shipyard workers at Titan. Image: G.Macgregor, CC rights.
The size of shipbuilding on the Clyde back in the glory days is today unimaginable. Several industrial cities in the UK had shipyards, Merseyside included; but just as Liverpool eclipsed all other cities as a port in the Victorian period, so Glasgow and the Clyde eclipsed all others in shipbuilding during the same era, producing 30,000 ships during 19th and 20th centuries. This equated to a third of all ships in the entire world, more than all the shipyards in Britain combined. It's a staggering statistic and earned Glasgow the title of "shipbuilding capital of the world".

I recently visited the former site of John Brown & Company Shipbuilders. The builder of choice for Cunard Line, John Brown was one of the 40 shipyards that prospered on the Clyde and produced some of the most famous vessels the world has ever seen. The Queen Mary, Queen Elizabeth, the Lusitania, the Aquitania, the Britannia, HMS Hood, the QE2 – they were all built there. And although it's been turned into an excellent tourist attraction with the help of EU funding, it remains a tragically haunting place. 100 hectares of open space. The largest of the slipways remain, upon which the QE2 would have been launched into the Clyde for fitting out. The classic Titan cantilever crane has been restored too, giving a tiny glimpse of the scale and sheer majesty of the vessels being built at the yard. But it was in the final years of the 'good times' at the shipyards that Jimmy Reid grew up.

RMS Queen Mary at Long Beach, California, now serving as a museum and hotel. Image: WPPilot, Wikimedia Commons, CC rights.
Reid was many things in his time, including a trade unionist, communist, Labour Party member, journalist and Rector of the University of Glasgow; however, he rose to international prominence in the early 1970s leading the Upper Clyde Shipbuilders (UCS) 'work-in' at Govan, Glasgow. The good times were over for British heavy industry, and this included marine engineering on the Clyde too. Increasing competition from abroad and a lack of investment meant that the yards increasingly required a state subsidy to complete their orders. UCS had gone into receivership and the Conservative government (led by Edward Heath) had decided that UCS - itself an amalgamation of five major Clyde shipyards several years earlier - should operate without state subsidy. The removal of these subsidies would immediately extinguish at least 6,000 jobs. Rather than adopt traditional forms of industrial action (e.g. strikes, sit-ins, etc.), the union leadership - spearheaded by Reid - decided to stage a 'work-in'. The union leadership were determined to complete the existing order book of ships and complete them to their traditional high levels of craftsmanship. Only this way would they dispel the myth of the 'work-shy' shipbuilder; only this way would they demonstrate their superior work ethic, project the best image to the British public, and demonstrate the viability of the yards. Said Reid famously, when addressing the shipyard workers:
"We are not going to strike. We are not even having a sit-in strike. Nobody and nothing will come in and nothing will go out without our permission. And there will be no hooliganism, there will be no vandalism, there will be no bevvying because the world is watching us, and it is our responsibility to conduct ourselves with responsibility, and with dignity, and with maturity."
This unique industrial action had integrity and was successful in the short-term, attracting international attention, sympathy and financial support (most notably from John Lennon and Yoko Ono). It is also why BAE Systems today have two ex-UCS yards on the Clyde, currently building the brand new high-end Type 45 Destroyers for the Royal Navy. Ultimately though, shipbuilding in Glasgow and on the Clyde today is a shade of its former self.
Titan cantilever crane at the former John Brown shipyard, Clydebank. Image: G.Macgregor, CC. rights
Jimmy Reid was born into a shipbuilding family. He left school at the age of 14 and pursued an apprenticeship in shipyard engineering. Despite leaving school with the minimum education possible, Reid went on to become one of the most talented political figures, orators, political thinkers and political leaders in British politics. He was fascinated by politics, reading about it constantly and studying political texts in the local public library almost every day. Sir Alex Ferguson – always astonished at Reid's knowledge - said at Reid's funeral:
"Our education was football, his education was the Govan library. He was never out of there."
Public libraries are often considered "the people's university". When Reid was made Rector of the University of Glasgow, "pompous" academics would ask which university he attended, to which Reid would reply: "Govan Library"! The public library moulded Reid, provided an education like no other and helped him develop an intellect which Sir Michael Parkinson described as "formidable". Reid personified the public library mission of education and lifelong learning available to all, regardless of age, skill level, or ability to pay. I have many issues with the running of public libraries today (some of which I might discuss in a future blog), but their importance in creating people like Jimmy Reid across Britain, and elsewhere in the world, can never be forgotten. And I'm not necessarily talking about people of Reid's political persuasion; but the importance of providing fantastic opportunities for education, enlightenment and betterment – and escape. The wealth of the Govan area during its shipbuilding heyday is reflected in its public buildings. The magnificent Town Hall, for example, with its neo-classical decoration, overlooking the quasi-futuristic architecture of Glasgow's redeveloped waterfront (and now used as a recording HQ for Franz Ferdinand). And Govan Public Library (or Elder Park Library as it is officially known) is no exception; a beautiful Victorian listed building situated within Elder Park. What a great location to build a public library; inviting local residents to escape the noise and dirt of shipbuilding and marine engineering to enjoy its salubrious surroundings and architectural splendour while reading an improving book! The Victorians had style - and they recognised the importance of the public library as an institution.
Elder Park Library, front elevation. Image: mike.thomson75.

It is a dangerous time for public libraries. Because they are such established institutions, people often take for granted that they will always exist, come rain or shine. Yet, the role of public libraries as the people's university remains as important as ever, not only to promote social inclusion and enable vulnerable people in society to engage with civil society, but providing opportunities for lifelong learning in dire economic circumstances. If some should go, where will the Jimmy Reids of tomorrow go? Some will argue that the increased penetration of broadband (circa 65%) makes some public libraries dispensable, but what about the remaining 35%, or those that are suddenly made redundant and can no longer afford their broadband bill? Where do they go if they want to develop knowledge in a particular discipline, or learn about IT? And just because someone has broadband does not mean that they can get access to all the information they might require (bibliographic databases?), or that the information will be any good. The Telegraph reckons Jimmy Reid's life would make a great biopic. I agree, so long as the splendour of Elder Park Library is maintained and we have plenty clips of Jimmy perusing the book stacks.

Thursday, 19 August 2010

Where are the WarGames students?

This morning I happened to enjoy "What's the point of ..." on BBC Radio 4. Motivated by an imminent spending review and inevitable cuts – as well as recent celebrations concerning the Battle of Britain – Quentin Letts put the RAF under some scrutiny. What's the point of the RAF in 2010? The issues and debates are well rehearsed, e.g. Are über futuristic fighter jets required for the conflicts the UK is likely to be engaged in? The outcome of the debate was inconclusive because few people can predict the types of conflicts that may emerge in the future.

One type of warfare which all commentators agreed was potentially imminent is cyber warfare. Not only is such warfare potentially imminent, but the UK (along with other NATO allies) is completely unprepared for a sophisticated or sustained attack. According to the programme only 24 people at the MoD are actively working on cyber security(!). Commentators agreed that funding had to be diverted from other armed services (i.e. RAF) to invest in cyber security. This means more advanced computing and information professionals to improve cyber security, but also to operate un-manned drones, manipulate intelligence data, and so forth. 'More Bill Gates-type recruits and fewer soldiers' was the message.

Although it wasn't given treatment in the programme, the conundrum for our cyber security – as well as our economy - is the declining numbers of students seeking to study computing science and information science at undergraduate/postgraduate level. This is a decline which is reflected more generally in the lack of school leaver interest in science and technology, something which – unless you have been living in a cave – the last Labour government and the current coalition are attempting to address. With the release of A-level results today and the massive demand for university places this year, some universities have been boosted by government grants designed to recruit extra students in science and technology. The coalition, in particular, sees it as a way of improving economic growth prospects; but it seems that the need to reverse this trend has become even more urgent given that we only have a small mini-bus full of 'cyber soldiers' – and, let's face it, five are probably on part-time contracts, two will be on maternity leave and one will be on long term sick leave.

It's a far cry from the 1980s Hollywood classic, 'WarGames' (1983). WarGames follows a young hacker (Matthew Broderick) who inadvertently accesses a US military supercomputer programmed to predict possible outcomes of nuclear war. Taking advantage of the unbelievably simple command language interface ("Can we play a global thermonuclear simulation game?" types Broderick) and Artificial Intelligence (AI) light years ahead 2010 state of the art, Broderick manages to initiate a nuclear war simulation believing it to be an innocent computer game. Of course, Broderick's shenanigans cause US military panic and almost cause World War III. I remember going to the petrol station with my father to rent WarGames on VHS as soon as it was available (yes – in the early 1980s petrol stations were often the place to go for video rentals! I suppose the video revolution was just kicking off...) and being thoroughly inspired by its depiction of computing and hacking. I wanted to be a hacker and was lucky enough to receive an Atari 800XL that Christmas, although programming soon gave way to gaming. Pac-Man anyone? Missile Command was pretty good too...

Monday, 26 July 2010

The end of social networking or just the beginning?

Today the Guardian's Digital Content blog carries an article by Charles Arthur in which we waxes lyrical about the fact that social networking - as a technological and social phenomenon - has reached its apex. As Arthur writes:
"I don't think anyone is going to build a social network from scratch whose only purpose is to connect people. We've got Facebook (personal), LinkedIn (business) and Twitter (SMS-length for mobile)."
Huh. Maybe he's right? The monopolisation of the social networking market is rather unfortunate and, I suppose, rather unhealthy - but it is probably and ultimately necessary owing to the current business models of social media (i.e. you've got to have a gargantuan user base to turn a profit). The 'big three' (above) have already trampled over the others to get to the top out of necessity.

However, Arthur's suggestion is that 'standalone' social networking websites are dead, rather than social networking itself. Social networking will, of course, continue; but it will be subsumed into other services as part of a package. How successful these will be is anyone's guess. This situation is contrary to what many commentators forecast several years ago. Commentators predicted an array of competing social networks, some highly specialised and catering for niche interests. Some have already been and gone; some continue to limp on, slowly burning the cash of venture capitalists. Researchers also hoped - and continue to hope - for open applications making greater use of machine readable data on foaf:persons using, erm, FOAF.

The bottom line is that it's simply too difficult to move social networks. For a variety reasons, Identi.ca is generally acknowledged to be an improvement on Twitter, offering greater functionality and open-source credentials (FOAF support anyone?); but persuading people to move is almost impossible. Moving results in a loss of social capital and users' labour, hence recent work in metadata standards to export your social networking capital. Yet, it is not in the interests of most social networks to make users' data portable. Monopolies are therefore always bound to emerge.

But is privacy the elephant in the room? Arthur's article omits the privacy furore which has pervaded Facebook in recent months. German data protection officials have launched a legal assault on Facebook for accessing and saving the personal data of people who don't even use the network, for example. And I would include myself in the group of people one step away from deleting his Facebook account. Enter diaspora: diaspora (what a great name for a social network!) is a "privacy aware, personally controlled, do-it-all, open source social network". The diaspora team vision is very exciting and inspirational. These are, after all, a bunch of NYU graduates with an average age of 20.5 and ace computer hacking skills. Scheduled for a September 2010 launch, diaspora will be a piece of open-source personal web server software designed to enable a distributed and decentralised alternative to services such as Facebook. Nice. So, contrary to Arthur's article, there are a new, innovative, standalone social networks emerging and being built from scratch. diaspora has immense momentum and taps into the increasing suspicion that users have of corporations like Facebook, Google and others.

Sadly, despite the exciting potential of diaspora, I fear they are too late. Users are concerned about privacy. It is a misconception to think that they aren't; but valuing privacy over social capital is a difficult choice for people that lead a virtual existence. Jettison five years of photos, comments, friendships, etc. or tolerate the privacy indiscretions of Facebook (or other social networks)? That's the question that users ask themselves. It again comes down to data portability and the transfer of social capital and/or user labour. diaspora will, I am sure, support many of the standards to make data portability possible, but will Facebook make it possible to output and export your data to diaspora? Probably not. I nevertheless watch the progress of diaspora closely and I hope, just hope they can make it a success. Good luck, chaps!

Monday, 19 July 2010

Google finally gets serious about the Semantic Web?

Google has been flirting with the Semantic Web recently, and we've talked about it occasionally on this blog. However, compared with other web search engines (e.g. Yahoo!) and the state of Semantic Web activity generally, Google has been slow to dive in completely. They have restricted themselves to rich snippets, using bits of RDFa and microformats, and making up their own too. Perhaps this was because their intention was always to purchase a prominent Semantic Web start-up company instead of putting in the spade work themselves? Perhaps so.

Google has this week announced the purchase of Metaweb Technologies. None the wiser?! Metaweb is perhaps most known for providing the Semantic Web community with Freebase. Freebase cropped up last year on this blog when we discussed the emergence of Common Tags. Freebase essentially represents a not insignificant hub in the rapidly expanding Linked Data cloud, providing RDF data on 12 million entities with URIs linking to other linked and Semantic Web datasets, e.g. DBpedia.

My comments are limited to the above; just thought this was probably an extremely important development and one to watch. A high level of social proof appears to be required before some tech firms or organisations will embrace the Semantic Web. But what greater social proof than Google? Google also appear committed to the Freebase ethos:
"[We] plan to maintain Freebase as a free and open database for the world. Better yet, we plan to contribute to and further develop Freebase and would be delighted if other web companies use and contribute to the data. We believe that by improving Freebase, it will be a tremendous resource to make the web richer for everyone. And to the extent the web becomes a better place, this is good for webmasters and good for users."
Very significant stuff indeed.

Tuesday, 13 July 2010

iStrain?

Usability guru Jakob Nielsen published details of a brief (but interesting) usability study on his Alertbox website last week. Nielsen was interested in exploring the differences that might exist between people reading long-form text on tablets and other devices. To be clear, this wasn't about testing the usability of devices per se; more about 'readability'.

Nielsen's research motivation was clear: e-book readers and tablets are finally growing in popularity and they are likely to become an important means of engaging in long-form reading in the future. However, such devices will only succeed if they are better than reading from PC or laptop screens and - the mother of all reading devices - the printed book. Nielsen and his assistants therefore performed a readability study of tablets, including the Apple's iPad and Amazon's Kindle, and compared these with books. You can read the article in full in your own time. It's a brief read at circa 1000 words. Essentially, Nielsen's key findings were that reading from a book is significantly quicker than reading from tablet devices. Reading from the iPad was found to be 6.2% slower and the Kindle 10.7% slower.

I recall ebook readers emerging in the late 1990s. At that time ebook readers were mysterious but exciting devices. After some above average ebook sales for a Stephen King best seller in 1999/2000 (I think), it was predicted that ebook readers would take over the publishing industry. But they didn't. The reasons for this were/are complex but pertain to a variety of factors including conflicting technologies, lack of interoperability, poor usability and so forth. There were additional issues, many of which some of my ex-colleagues investigated with their EBONI project. One of the biggest factors inhibiting their proliferation was the issue of eye strain. The screens on early ebook readers lacked sufficient resolution and were simply small computer screens which came with the associated eye strain issues for long-form reading, e.g. glare, soreness of the eyes, headaches, etc. Long-form reading was simply too unpleasant; which is why the emergence of the Kindle, with its use of e-ink, was revelatory. The Kindle - and readers like it - have been able to simulate the printed word such that eye strain issues are no longer an issue.

Nielsen has already attracted criticism regarding flaws in his methodology; but in his defence he did not purport his study to be rigorously scientific, nor has he sought publication of his research in the peer-reviewed research literature. He wrote up his research in 1000 words for his website, for goodness sake! In any case, his results were to be expected. Applegeeks will complain that he didn't use enough participants, although those familiar with the realities of academic research will know that 30-40 participant user studies are par for the course. However, there is one assumption in Nielsen's article which is problematic and which has evaded discussion: iStrain. Yes - it's a dreadful pun but it strikes at the heart of whether these devices are truly readable or not. Indeed, how conducive can a tablet or reader be for long-form reading if your retinas are bleeding after 50 minutes reading? Participants in Nielsen's experiment were reading for around 17 minutes. Says Nielsen:
"On average, the stories took 17 minutes and 20 seconds to read. This is obviously less time than people might spend reading a novel or a college textbook, but it's much longer than the abrupt reading that characterizes Web browsing. Asking users to read 17 minutes or more is enough to get them immersed in the story. It's also representative for many other formats of interest, such as whitepapers and reports."
...All of which is true, sort of. But in order to assess long-form reading participants need to be reading for a lot, lot longer than 17 minutes, and whilst the iPad enjoys a high screen resolution and high levels of user satisfaction, how conducive can it be to long-form reading? And herein lies a problem. The iPad was never really designed as an e-reader. It is a multi-purpose mobile device which technology commentators - contrary to all HCI usability and ergonomics research - seem to think is ideally suited to long-form reading. It may rejuvenate the newspaper industry since this form of consumption is similar to that explained above
by Nielsen, but the iPad is ultimately no different to the failed e-reading technologies of ten years ago. In fact, some might say it is worse. I mean, would you want to read P.G. Wodehouse through smudged fingerprints?! The results of Nielsen's study are therefore interesting but they could have been more informative had participants been reading for longer. A follow up study is order of the day and would be ideally suited to an MSc dissertation. Any student takers?!

(Image (eye): Vernhart, Flickr, Creative Commons Attribution-NonCommercial-ShareAlike 2.0 Generic)
(Image (beware): florian.b, Flickr, Attribution-NonCommercial 2.0 Generic)

Monday, 28 June 2010

Love me, connect me, Facebook me

This is just a quick post to alert readers to the recent existence of a dedicated MA/MSc/PG Dip. Information & Library Management Facebook group page. The page is for current and prospective ILM students, but should also prove useful to alumni and enable networking between former students and/or other information professionals.

Being a Facebook group page it is accessible to everyone; however, those of you with Facebook accounts can become fans to be kept abreast of programme news, events, research activity, industry developments and so forth. Click the "Like it!" button!

Wednesday, 23 June 2010

Visualising the metadata universe

No blog postings for almost three months and then two come along at once...  I thought it would be worth drawing to the attention of readers the recent work of Jenn Riley of Indiana University.  Jenn is currently metadata guru for the Indiana University Digital Library Program and yesterday on the Dublin Core list she announced the output of a project to build a conceptual model of the 'metadata universe'. 

As evidenced by some of my blogs, there are literally hundreds of metadata standards and structured data formats available, all with their own acronym.  This seems to have become more complicated with the emergence of numerous XML based standards in the early to mid noughties, and the more recent proliferation of RDF vocabularies for the Semantic Web and the associated Linked Data drive.  What formats exists?  How do they relate to each other?  For which communities of practice are they optimised, e.g. information industry or cultural sector?  What are the metadata, technical standards, vocabularies that I should be congnisant of in my area?  And so the question list goes on...


These questions can be difficult to answer, and it is for this reason that Jenn Riley has produced a gigantic poster diagram (above) entitled, 'Seeing standards: a visualization of the metadata universe'.  The diagram achieves what a good model should, i.e. simplifying complex phenomena and presenting a large volume of information in a condensed way.  As the website blurb states:
"Each of the 105 standards listed here is evaluated on its strength of application to defined categories in each of four axes: community, domain, function, and purpose. The strength of a standard in a given category is determined by a mixture of its adoption in that category, its design intent, and its overall appropriateness for use in that category."
A useful conceptual tool for academics, practitioners and students alike.  A glossary of metadata standards in either poster or pamphlet form is also available.