Wednesday, 27 May 2009

Image searching with Creative Commons

Student information literacy skills have been discussed on the blog before. In short, they are woeful. One area where students tend to have little understanding is in the area of intellectual property rights (IPR). The situation might be looking better for digital music, but in my experience it remains poor for other digital artefacts, particularly images. 'Twas only a few weeks ago while was in a lab with some undergraduate students for a web technologies module when I discovered most of them were ripping images from the web for inclusion within their information gateways. While this can (in some circumstances) be tolerated within the confines of an educational institution, it remains copyright infringement owing to copying by 'reprographic means' - and this isn't behaviour we want to become habitual in our graduates. My brother (a graphic designer and new media guru) has spun me many a yarn about ex-colleagues who have been shown their P45 for engaging in IPR theft (e.g. reusing someone's basic design or photograph).

All of this is veering away from the original reason for this blog though, which is to draw attention to some new image searching functionality on Yahoo! Image Search. Following on nicely from the Search Options post, the Yahoo! Search Blog has just announced the inclusion of some extra search filters for image result sets. Not only is it better than Google (and more accurate?), but it also includes a useful Creative Commons (CC) filter. Using a similar interface to Yahoo! Search Assist, Yahoo! Image Search allows users to apply a CC checkbox to filter for images, with specific filters included for commercial reuse and/or remixing. This is particularly useful to embellish those PowerPoint presentations or to illustrate a blog, or for those undergraduate students building an information gateway, or to avoid getting a P45!

There appears to be a downside, unfortunately. When I saw the Yahoo! Search Blog announcement I thought (perhaps naively) that Yahoo! was starting to put into practice its commitment to metadata, Semantic Web specifications, and other structured data. Since I know my personal homepage is indexed by Yahoo! and uses XHTML+RDFa to notify intelligent agents that its page content falls under a Creative Commons Attribution 3.0 License, I thought I'd put an Image Search to the test. Providing the CC namespace is referenced, the XHTML+RDFa required is simple. For example:

<p>Content on <a href="http://www.staff.ljmu.ac.uk/bsngmacg/" property="cc:attributionName" rel="cc:attributionURL">George Macgregor</a>'s website is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by/3.0/">Creative Commons Attribution 3.0 License</a></p>

...and with specific CC reference to my foaf:depiction...

<img src="img/georgedepiction.jpg" alt="Image of George Macgregor" rel="license" href="http://creativecommons.org/licenses/by/3.0/" property="foaf:depiction" content="George Macgregor"/>

My filtered CC search was unsuccessful though. This disappointed me; but then I observed the following notice:
"Note: Only Flickr images are supported currently."
Flickr – which is a subsidiary of Yahoo! – has allowed users to conduct advanced searches of its publicly uploaded images for quite some time. This has included CC searching. And it would appear that Yahoo! has integrated Flickr searching functionality into Image Search, albeit with some nice tweaks. If I had read their blog in its entirety I would have realised this; I just clicked the link such was my excitement about Yahoo! Image Search!

It's useful to have this functionality within a conventional searching tool, but it is disappointing that Image Search isn't using cleverer means of doing it (e.g. RDFa) rather than relying on the preferences of Flickr users when they upload their images. Don't get me wrong, this is useful and most welcome, and it will save me time on occasion, but it would be exciting to crack CC image searching beyond the controlled Flickr environment. Hopefully the 'currently' in "Only Flickr images are supported currently" will mean that my expectations will be met soon…

Monday, 18 May 2009

Light relief: Celtic fringes erased by WolframAlpha?!

With WolframAlpha launched on Friday, I spent much of my weekend trying to get a 'computable request' to compute. Not until Monday morning did a request compute – but its performance has been getting better ever since so hopefully we will all have more time to experiment with it over coming days and weeks...

Like me, Gwenda Mynott has been testing WolframAlpha and has been searching for things that, a) you have a good knowledge of, and, b) a topic that WolframAlpha can easily compute. Places are good for this (e.g. countries, towns, cities, etc.), and Stephen Wolfram computes multiple locations to good effect in his demonstrations; however, Gwenda tried to 'compute' Wales and arrived at some bizarre results. Check them out. WolframAlpha doesn't retrieve data pertaining to the constituent nation of the United Kingdom of Great Britain and Northern Ireland (i.e. Wales as you or I would tend to know it!), but a small town in South Yorkshire by the name of Wales (?) The only other obvious option WolframAlpha provides is Wales (New York, USA), which is equally amiss.

Hmmmmm. If this is the result for Wales, what are the results for the rest of the UK? Well, that's equally controversial. England appears to be synonymous with the United Kingdom of Great Britain and Northern Ireland. Scotland is referred by WolframAlpha back to the Kingdom of Scotland, which ceased to exist after the Act of Union in 1707. Worse than that, Northern Ireland doesn't even exist! ("WolframAlpha isn't sure what to do with your input") Cornish nationalists will also be dismayed to learn that Cornwall (Canada) is the only one that counts.

Is this a systematic attempt to erase the history, culture and memory of the Celtic fringes?! Of course not. The results might be strange, but from a knowledge engine point of view – and ontologically speaking - Wales, Scotland and Northern Ireland are subsumed by the larger geographical and political entity of the UK, so it's understandable that WolframAlpha computes the answer in this way. Still, the England/UK synonymy is a bit odd and must have been encoded by someone somewhere sometime!

Experiment away, folks - and I would encourage everyone to post their most bizarre / illogical data results as comments to this blog. A prize will go to the most outlandish!

Friday, 15 May 2009

Some more 'Search Options'...

I promised not to blog about Google any time soon for fear the blog becomes known as the unofficial Google blog. After some consideration I thought, 'pish posh!' Anyway, the post has a wider remit than just Google...honest!

The absence of retrieval aids for Google users (oh no, not again - I hear you cry!) has been discussed at great length on this blog before. To appreciate the extent of this deficiency we need only peruse some innovative rival search engines such as Ask (recently re-branded back to Ask Jeeves), Yahoo!, or Clusty. Google has been making changes though and today the Official Google Blog announced some further enhancements to the universal Google search interface. Simply called, 'Search Options', these tools let you "slice and dice" results, apply rudimentary filters, and generate alternative views of results. Search Options does a little bit more to help the user in query formulation (the area where I think Google is weakest), but also offers some useful functionality once you have your results.



Check out our usual canned search for 'communism in Russia'; 'click' the 'Search options…' link in the top left had corner of the interface to reveal the Search Option tools.

Filters are available for videos, forums and reviews (the latter being fairly useful if you are shopping). Various publication time filters are also available. Nothing here is particularly mind blowing though.

Search Options gets a bit more interesting when the search display options are explored in a little detail. Firstly, it's possible to request details of related searches. These are displayed in a better page location than before and look similar to Yahoo! Search Assist. But it is now also possible to select the 'Wonder Wheel' which generates a visualisation of the related terms. I'm unsure how useful the Wonder Wheel really is, particularly as the true nature of the relationships between terms is impossible for Google to represent other than in syntactic terms; this is something the Semantic Web community is obviously trying to resolve.

Most interesting though is the 'Timeline' tool. This allows results to be displayed along, erm, yup, a timeline. The timeline is clickable allowing the user to drill down into particular temporal zones and to view resources relating to that zone. I use the word 'interesting' because although the timeline is probably quite useful for historical research, its moment of introduction is the most interesting part. Indeed, the timeline functionality looks in part like Google is bracing itself for the release of WolframAlpha, which is due any day now (or tonight?) – and I wouldn't be at all surprised if this announcement was an attempt to steal some of its thunder. This appears to have been combined with the demonstration of Google Squared at the Google Searchology conference a few days ago. No Google Squared prototypes appear to be available for us to experiment with, but TechCrunch got a sneaky peak at Searchology (view the YouTube video below). Google Squared is, in essence, Google's answer to WolframAlpha.

For me the most interesting news to emerge alongside Search Options is Google's desire to make greater use of RDFa. RDFa is probably a little pedestrian for me, but it's better than nothing – and at least there is a clear intention of using some Semantic Web specifications. It's just a shame Yahoo! announced something similar but more radical almost 18 months ago.

Friday, 1 May 2009

LCSH as Linked Data ... officially!

Yesterday was, in my estimation, pretty historic. The Library of Congress officially launched the LC Authorities and Vocabularies service. You might recall a previous post relating to lcsh.info in which I lamented the LC's decision to pull down a SKOS demonstrator of LCSH, explicitly designed to explore the possibilities of Linked Data and dereferenceable URIs. All the background is in the previous post; but the whole episode appears to have been a PR disaster for LC.

The great news is that the LC Authorities and Vocabularies service (let's call it LCAV henceforth, shall we?) officially re-launched lcsh.info in a bigger, better and much improved form. The service essentially enables both humans and machines to access a plethora of LC authority data. Like lcsh.info, the service employs Semantic Web approaches to exposing this data and implements approaches to Linked Data by exposing and linking data on the Web via dereferenceable URIs.

Five minutes exploring the website reveals that LCAV serves up the entire LCSH for free, with incredible search and browse functionality, leaving Connexion in the shade. The concept URIs point to detailed data modelled in SKOS as RDFa for human readability, but with links to SKOS as RDF/XML, N-Triples and (the less familiar?) JSON for machine processing. RDF graphs can even be visualised by clicking, well, the 'visualize' tab – incredible. Mappings to other vocabularies are also provided.
On top of all this, LCSH can be downloaded in its entirety as RDF/XML or N-Triples (SKOS)! LCAV also indicate that further authority data will be made available soon.

Make no bones about it, this is historic stuff, not only because the service is so good but because this terminological data is no longer locked down. I think it's important to stroke our imaginary beards over the significance of the LC's change of direction. Is this the beginning of the end for locked down terminological data?! Will they be like dominoes henceforth? A fiver says DDC does the same by the end of the year. Any takers???

Thursday, 30 April 2009

WolframAlpha and destructive hype

If you have been plugged into the search engine or technology news feeds over recent months you may have encountered the excitement surrounding WolframAlpha. WolframAlpha is a Web search tool to be launched in May which – apart from having a good name – is anticipated to be the next Google exterminator. Although being touted as a destroyer of Google, the technology commentators indicate that WolframAlpha will inhabit an entirely different intellectual space on the Web.

WolframAlpha is described by its creators as a "computational knowledge engine" which, instead of retrieving resources using conventional automatic indexing methods, dynamically computes the answers to a wide variety of questions. The way in which it does this remains a mystery, but we do know that it models particular areas of knowledge. It then combines this with a vast repository of curated data harvested from disparate data sources and some ingenious natural language processing algorithms to represent knowledge. These knowledge representations can then be queried to answer real questions. Stills sounds like an enigma; but it must work on some level given the hype around it. Mustn't it?!

The brainchild of Dr. Stephen Wolfram (purveyor of computer algebra), WolframAlpha has had information and computer scientists and technology commentators salivating for months. The trouble is that while the incessant hype continues, an increasing number of people (me, but some commentators) are growing increasingly cynical of its true capabilities; we want to see a demo, or some kind of prototype. Mindful that cynicism could be spreading, Wolfram unveiled his creation yesterday for the first time at the Harvard University Berkman Center for Internet & Society (via a sold-out Web cast - clip from YouTube below). This demonstration appears to have further stimulated the hype (judging by some headlines), but has simultaneously added to the increasing cynicism. Hype and 'vapourware' exasperates people. And this is where the hype could actually be death of WolframAlpha, rather than Google.

In reality, it certainly sounds like WolframAlpha is not out to compete with Google; but it doesn't matter, this is how it is being described in the media and WolframAlpha hasn't tried to dispel the myth. In his blog, Wolfram describes WolframAlpha as "a new paradigm for using computing and the Web". It immediately provides people with a Google yardstick and false expectations; most new users will not understand that WolframAlpha is an entirely different beast. But more importantly, it's setting WolframAlpha up for an almighty fall.

Remember Cuil? People also thought Cuil was going to change the face of searching but it failed. It was hotly anticipated and was hyped, arguably more, than WolframAlpha. This hype did it no favours when it crashed on its launch day. It's only been 9 months since Cuil was officially launched, yet we never hear about it, nor do any of us use it. In part, this is because its indexes are so poor. My LJMU profile page was updated on 08 November 2008, almost 6 months ago; yet, Cuil still returns this page as it was on 07 November 2008 as a result. This is extremely feeble when you consider that Microsoft Live Search refreshes its indexes every 20 days.

Along with the hype, this was Cuil's 'blind spot'. A blind spot is normally tolerated in the early days of an innovative Web tool, but inflated expectations breeds intolerance. WolframAlpha is bound to have its own blind spot; what will it be and will users be tolerant until it's fixed? Probably not. They therefore have to get it right on launch day.

The moral of this tale is simple. Hyperbole must end. It's destructive and in the long run it does nobody any favours.

Tuesday, 7 April 2009

Web 2.0? Show me the money!

Just a quick post... Today the Guardian blog reports on the financial woes of YouTube. I don't suppose we should be particularly surprised to learn that according to some news sources YouTube is due to drop $470 million this year. When this figure is compared to the $1.65 billion pricetag Google paid a couple of years ago we can appreciate the magnitude of their YouTube predicament. The majority of this loss is attributable to the failure of advertising to bring home the bacon; a recurring issue on this blog. But huge running costs, copyright and royalty issues have played their part too. Google is reportedly interested in purchasing Twitter, but surely their failure to monetise YouTube - a service arguably more monetiseable (?) than Twitter - should have the alarm bells ringing at Google HQ?

I find the current crossroads for many of these services utterly fascinating. I don't have any solutions for any of these ventures, other than to make sure you have a business model before starting any business. Would RBS give me a business loan without a business plan and a robust revenue model? Probably not. But then they are not giving loans out these days anyway...

Wednesday, 25 March 2009

ASK conundrum revisited...again!

I posted a blog about Google's eye tracking research last month. I'm loathed to discuss Google again lest the ISG blog becomes known as the unofficial Google blog; however, the latest post on the Official Google Blog is worthy of some comment...

You might recall another post I made regarding search engine research and development, particularly in the area of information retrieval (IR) aids for users. In this posting I summarised Belkin's research and theories regarding the Anomalous State of Knowledge (ASK). Most of this and subsequent research has sought to introduce IR aids for the user so that they can better solve their ASK conundrum. This assistance varies but often takes the form of query expansion (in its various permutations), browsable subject trees to stimulate query formulation, relevance feedback, and so forth. Providing such tools in systems based on automatic indexing is difficult, but we noted that some search engines have introduced some effective retrieval aids, all designed to alleviate the ASK problem. For example, Yahoo! provides its search assist tool, Clusty provides related concept clusters, and Ask provides other similar tools. Their accuracy in IR varies widely, but overall they prove useful to user. Unfortunately, we also noted that Google provides few user aids comparable to those above, arguably relying more on its PageRank algorithm. Not any longer...

Today Google launched some interface functionality not dissimilar to Yahoo! search assist and Clusty. Their assistance provides some suggested related searches and some extra result summary text for particular results. Receiving this assistance depends on the nature of your query, so have a look at this canned search: 'communism in Russia'. This isn't bad and is better than nothing; but does it really measure up to the aids provided by competing search engines? Compare the results for these canned searches and the IR aids provided for the user by the systems we've discussed already:
Google's attempts appear quite pedestrian by comparison. Yahoo! and Clusty, for example, make their aids readily available so that the user can affect changes in their information seeking behaviour, but Google's tools are far less visible, less detailed, and offer far less functionality. Since a lot of research indicates that many users will not scroll below the 'golden triangle' (i.e. to the bottom of the first result set), it is entirely feasible to think that these 'related search' aids will go unnoticed by the disoriented information seeker.

It is good to see Google deploying user query aids and reacting to developments in other IR systems, but it appears that it will be some time before Google can be said to alleviate users' Anomalous State of Knowledge.

Tuesday, 17 March 2009

Who's going to teach our "stuff"? and who's going to learn it?

Is this the right forum for this? We all know of the imminent and unwelcome restructuring facing us. Where does the future lie for this discipline, or can we even define what our discipline is? We're constantly reminded now that our HE degrees are products, our students are customers, and so what are we, retailers? Compared to many other "products" in this HE marketplace our products are relatively unpopular despite the fact that there are fewer universities providing what we do compared to a decade ago.
I constantly struggle to explain what it is we do and can therefore understand why students have difficulty in placing it in context. Perhaps a business school isn't the right environment but then neither is a computing department, or it doesn't seem to be, and we don't fit into education or anywhere else.
What is the long term future for this discipline, whatever it is?

It's St. Patrick's day (not St. Paddy's, or St. Pat's or Paddy's day) so I'm off for Guinness in the local.

Thursday, 19 February 2009

Text is the new GUI?

We've got a software student ( from another University) working on his final year project at my day job. He is busy adding a 'speech' interface to our Laboratory tracking system. The idea is that while the scientist has their hands in the fume cupboard they don't want to be messing with a mouse or a keyboard why not engage with the computer by voice. This is after all how they did it in the old future in the movies.

Alas the student has got a bit sidetracked into the excitement of speech recognition and synthesis, in an attempt to get him to some sort of conclusion of his project I have suggested that most of the academic benefit of the project could be got by just having a text input and output (although useless in the fume cupboard). Once you go there it starts you thinking about how we interact with systems by voice. We are of course now used to listening to SatNav and some folk order their phones to phone the wife.

While pondering this and Georges earlier post of Twitter Library fees, I was thinking about an article about how fans have put together Twitter accounts of their favourite T.V. characters so that we can see when they are having a sandwich during the week. It made me wonder whether we might soon be engaging with various systems through the power of text rather than super 3D graphical interfaces.

This will be a shame as much of the design thought in web design and business application design has been about mice and windows more or less. In Liverpool this has led to some success for the hybrid 'programmer/graphic designer', perhaps if we are going to deal in a flow of text it will be the Hybrid 'programmer/DJ' or at least 'programmer/Linguist' who lead the way.

We surround ourselves in an increasing sense of a flow of consciousness through Twitter/Facebook etc. Surely this is going to include a Twitter from machines.

"Your fridge is enjoying a quiet day."
"Your car is worrying that it's service is due this week."
"Your door notes that fido is standing at it and wants to go out."

This will lead us to a wish to push application outputs into twitter like streams for our apps to respond to our own twitters. Text (or speech) may be the new GUI. Of course we will have to know a lot more about parsing and extracting meaning and identity from these streams of conciousness.

Thursday, 12 February 2009

FOAF and political social graphs

While catching up on some blogs I follow, I noticed that the Semantic Web-ite Ivan Herman posted comments regarding the US Congress SpaceBook – a US political answer to Facebook. He, in turn, was commenting on a blog made by the ProgrammableWeb – the website dedicated to keeping us informed of the latest web services, mashups, and Web 2.0 APIs.

From a mashup perspective, SpaceBook is pretty incredible, incorporating (so far) 11 different Web APIs. However, for me SpaceBook is interesting because it makes use of semantic data provided via FOAF and the microformat, XFN. To do this SpaceBook makes good use of the Google Social Graph API, which aims to harness such data to generate social graphs. The Social Graph API has been available for almost a year but has had quite a low profile until now. Says the API website:
"Google Search helps make this information more accessible and useful. If you take away the documents, you're left with the connections between people. Information about the public connections between people is really useful -- as a user, you might want to see who else you're connected to, and as a developer of social applications, you can provide better features for your users if you know who their public friends are. There hasn't been a good way to access this information. The Social Graph API now makes information about the public connections between people on the Web, expressed by XFN and FOAF markup and other publicly declared connections, easily available and useful for developers."
Bravo! This creates some neat connections. Unfortunately – and as Ivan Herman regrettably notes - the generated FOAF data is inserted into Hilary Clinton’s page as a page comment, rather than as a separate .rdf file or as RDFa. The FOAF file is also a little limited, but it does include links to her Twitter account. More puzzling for me though is why the embedded XHTML metadata does not use Qualified Dublin Core! Let's crank up the interoperability, please!

Friday, 6 February 2009

Information seeking behaviour at Google: eye-tracking research

Anne Aula and Kerry Rodden have just published a posting on the Official Google Blog summarising some eye-tracking research they have been conducting on Google's 'Universal Search'. Both are active in information seeking behaviour and human-computer interaction research at Google and are well published within the related literature (e.g. JASIST, IPM, SIGIR, CHI, etc.).

The motivation behind their research was to evaluate the effect incorporation of thumbnail images and video within a research set has on user information seeking behaviour. Previous information retrieval eye-tracking research indicates that users scan results in order, scanning down their results until they reach a (potentially) relevant result, or until they decide to refine their search query or abandon the search. Aula and Rodden were concerned that the inclusion of thumbnail images might distract the "well-established order of result evaluation". Some comparative evaluation was therefore order of the day.
"We ran a series of eye-tracking studies where we compared how users scan the search results pages with and without thumbnail images. Our studies showed that the thumbnails did not strongly affect the order of scanning the results and seemed to make it easier for the participants to find the result they wanted."
A good finding for Google, of course; but most astonishing is the eye-tracking data. The speed with which users scanned result sets and the number of points on the interface they scanned was incredible. View the 'real time' clip below. A dot increasing in size denotes the length of time a user spent pausing at that specific point in the interface or result set. Some other interesting discoveries were made – the full posting is essential reading.

Friday, 30 January 2009

Wikipedia: the new Knol?

Like many people I use Wikipedia quite regularly to check random facts. The strange aspect of this behaviour is that once I find the relevant fact, I have to immediately verify its provenance by conducting subsequent searches in order to find corroborative sources. It makes one wonder why one would use it in the first place.

Wikipedia continues to be plagued by a series of high profile malicious edits. Unfortunately, many of these edits aren't necessarily malicious. They are just wrong or inaccurate. There are probably hundreds of thousands of inaccurate Wikipedia articles, perhaps just as many hosting malicious edits; but it takes high profile gaffs to affect real change. On the day of Barack Obama's inauguration, Wikipedia reported the deaths of West Virginia's Robert Byrd and Edward Kennedy, who had collapsed during the inaugural lunch. Both reports were false.

This event appears to have compelled Jimmy Wales into being more proactive in improving the accuracy and reliability of Wikipedia. Under his proposals many future changes to articles would need to be approved by a group of vetted editors before being published. For me this news is interesting, particularly as it emerges barely two weeks after Google announced that the 100,000th Knol had been created on their 'authoritative and credible' answer to Wikipedia: Knol. Six or seven months ago the discussions focussed on how Knol was the new Wikipedia, now it appears as if Wikipedia might become the new Knol. How bizarre is that?!

Tightening the editing rules of Wikipedia has been on the agenda before and in 2007 this blog discussed how the German Wikipedia was conducting experiments which saw only trusted Wikipedians verifying changes to articles. So, will tightening the editing of Wikipedia make it the new Knol? The short answer is 'no'. Some existing Wikipedia editors can already exert authoritarian control over particular articles and can – in some cases – give the impression that they too have an axe to grind on particular topics. 'Wikipinochets' anyone? Moreover, Knol benefits from its "moderated collaboration" approach, with Knols being created by subject experts whose credentials have been verified. Wikipedia isn't going anywhere near this. Guardian columnist, Marcel Berlins, is probably right about Wikipedia when he states:
"I don't think there's a way of telling what proportion of Wikipedia entries are deficient, whether because of the writer's bias, mischief or lack of knowledge. It's clear that a significant number are questionable, sufficient to lead us to suspect all entries. But to do the right thing - vetting all contributors or contributions - would be impractical and hugely expensive. There is no easy solution. We many just have to accept that Wikipedia's undoubted usefulness comes at the price of occasional - perhaps frequent - inaccuracy. That is a sad conclusion to reach about an encyclopedia."
Oh well, back to verifying random facts found on Wikipedia.

Wednesday, 24 December 2008

SKOS-ifying Knowledge Organisation Systems: a continuing contradiction for the Semantic Web

A few days ago Ed Summers announced on his blog that he was shutting down lcsh.info. For those that don't know, lcsh.info was a Semantic Web demonstrator developed by Ed with the expressed purpose of illustrating how the Library of Congress Subject Headings (LCSH) could be represented and its structure harnessed using Simple Knowledge Organisation Systems (SKOS). In particular, Ed was keen to explore issues pertaining to Linked Data and the representation of concepts using URIs. He even hoped that the URIs used would be Cool URIs, linking eventually to a bona fide LCSH service were one ever to be released. Sadly, it was not to be... The reasons remain unclear but were presumably related to IPR. As the lcsh.info blog entry notes, Ed was compelled to remove it by the Library of Congress itself. The fact that he was the LC's resident Semantic Web buff probably didn't help matters, I'm sure.

SKOS falls within my area of interest and is an initiative of the Semantic Web Deployment Working Group. In brief, SKOS is an application of RDF and RDFS and is a series of evolving specifications and standards used to support the use of knowledge organisation systems (KOS) (e.g. information retrieval thesauri, classification schemes, subject heading systems, taxonomies or any other controlled vocabulary) within the framework of the Semantic Web. The Semantic Web is many things of course; but it is predicated upon the assumption that there exists communities of practices willing and able to create the necessary structured data (generally applications of RDF) to make it work. This might be metadata, or it might be an ontology, or it might be a KOS represented in SKOS. The resulting data can then be re-used, integrated, interconnected, queried and is open. When large communities of practice fail to contribute, the model breaks down.

There is a sense in which the Semantic Web has been designed to bring out the schizophrenic tendencies within some quarters of the LIS community. Whilst the majority of our community has embraced SKOS (and other related specifications), can appreciate the potential and actively contributes to the evolution of the standards, there is a small coterie that flirts with the technology whilst simultaneously shirking at the thought of exposing hitherto proprietary data. It's the 'lock down' versus 'openness' contradiction again.

In a previous research post I was involved with the High-Level Thesaurus (HILT) research project and continue my involvement in an consultative capacity. HILT continues to research and develop a terminology web service providing M2M access to a plethora of terminological data, including terminology mappings. Such terminological data can be incorporated into local systems to improve local searching functionality. Improvements might include, say, implementing a dynamic hierarchical subject browsing tree, or incorporating interactive query expansion techniques as part of the search interface, for example. An important - and the original motivation behind HILT - is to develop a 'terminology mapping server' capable of ameliorating the "limited terminological interoperability afforded between the federation of repositories, digital libraries and information services comprising the UK Joint Information Systems Committee (JISC) Information Environment" (Macgregor et al., 2007), thus enabling accurate federated subject-based information retrieval. This is a blog so detail will be avoided for now; but, in essence, HILT is an attempt to provide a terminology server in a mash-up context using open standards. To make the terminological data as usable as possible and to expose it to the Semantic Web, the data is modelled using SKOS.

But what happens to HILT when/if it becomes an operational service? Will its terminological innards be ripped out by the custodians of terminologies because they no longer want their data exposed, or will the ethos of the model be undermined as service administrators permit only HE institutions or charitable organisations from accessing the data? This isn't a concern for HILT yet; but it is one I anticipated several years ago. And the sad experience of lcsh.info illustrates that it's a very real concern.

Digital libraries, repositories and other information services have to decide where they want to be. This is a crossroads within a much bigger picture. Do they want their much needed data put to a good use on the Web, as some are doing (e.g. AGROVOC, GEMET, UKAT)? Or do they want alternative approaches to supplant them entirely (i.e. LCSH)? What's it gonna be, punks???

Monday, 8 December 2008

Wikipedia censorship: allusions to 'Smell the Glove'?

Another Wikipedia controversy rages, this time over censorship. Over the past two days, the Internet Watch Foundation informed some ISPs that an article pertaining to an album by the 'classic' German heavy metal band, Scorpions, may be illegal. Leaving aside the fact that Scorpions is one of many groups to have similar imagery on their record sleeves (the eponymous 1969 debut album by Blind Faith, Eric Clapton's supergroup, being another obvious example), am I the only person to notice the similarities with fictional rockumentary, This Is Spinal Tap?

Like metal, censorship is a heavy topic; but I thought this tenuous linkage with This Is Spinal Tap might be a welcome distraction from the usual blog postings, which are necessarily academic. Those of you familiar with said film might recall the controversy surrounding the proposed (tasteless) art work for Spinal Tap's new album (Smell the Glove), which in the end gets mothballed owing to its indecent nature. Getting into trouble over sleeve art is part and parcel of being in a heavy metal band it would seem! Enjoy the winter break, people!

Friday, 5 December 2008

Some general musings on tag clouds, resource discovery and pointless widgets...

The efficacy of collaborative tagging in information retrieval and resource discovery has undergone some discussion on this blog in the past. Despite emerging a good couple of years ago – and like many other Web 2.0 developments – collaborative tagging remains a topic of uncertainty; an area lacking sufficient evaluation and research. A creature of collaborative tagging which has similarly evaded adequate evaluation is the (seemingly ubiquitous!) 'tag cloud'. Invented by Flickr (Flickr tag cloud) and popularised by delicious (and aren't you glad they dropped the irritating full stops in their name and URL a few months ago?), tag clouds are everywhere; cluttering interfaces with their differently and irritatingly sized fonts.

Coincidentally, a series of tag cloud themed research papers were discussed at one of our recent ISG research group meetings. One of the papers under discussion (Sinclair & Cardew-Hall, 2008) conducted an experimental study comparing the usage and effectiveness of tag clouds with traditional search interface approaches to information retrieval. Their work is welcomed since it constitutes one of the few robust evaluations of tag clouds since they emerged several years ago.

One would hate to consider tag clouds as completely useless – and I have to admit to harbouring this thought. Fortunately, Sinclair and Cardew-Hall found tag clouds to be not entirely without merit. Whilst they are not conducive to precision retrieval and often conceal relevant resources, the authors found that users reported them useful for broad browsing and/or non-specific resource discovery. They were also found to be useful in characterising the subject nature of databases to be searched, thus aiding the information seeking process. The utility of tag clouds therefore remains confined to the search behaviours of inexperienced searchers and – as the authors conclude - cannot displace traditional search facilities or taxonomic browsing structures. As always, further research is required...

The only thing saving tag clouds from being completely useless is that they can occasionally assist you in finding something useful, perhaps serendipitously. What would be the point in having a tag cloud that didn't help you retrieve any information at all? Answer: There wouldn't be any point; but this doesn't stop some people. Recently we have witnessed the emergence of 'tag cloud generation' tools. Such tools generate tag clouds for Web pages, or text entered by the user. Wordle is one such example. They look nice and create interesting visualisations, but don't seem to do anything other than take a paragraph of text and increase the size of words based on frequency. (See the screen shot of a Wordle tag cloud for my home page research interests.)


OCLC have developed their very own tag cloud generator. Clearly, this widget has been created while developing their suite of nifty services, such as WorldCat, DeweyBrowser, FictionFinder, etc., so we must hold fire on the criticism. But unlike Wordle, this is something OCLC could make useful. For example, if I generate a tag cloud via this service, I expect to be able to click on a tag and immediately initiate a search on WorldCat, or a variety of OCLC services … or the Web generally! In line with good information retrieval practice, I also expect stopwords to be removed. In my example some of the largest tags are nonsense, such as "etc", "specifically", "use", etc. But I guess this is also a fundamental problem with tagging generally...

OCLC are also in a unique position in that they have access to numerous terminologies. This obviously cracks open the potential for cross-referencing tags with their terminological datasets so that only genuine controlled subject terms feature in the tag cloud, or productive linkages can be established between tags and controlled terms. This idea is almost as old as tagging itself but, again, has taken until recently to be investigated properly. Exploring the connections between tags and controlled vocabularies is something the EnTag project is exploring, a partner in which is OCLC. In particular, EnTag (Enhanced Tagging for Discovery) is exploring whether tag data, normally typified by its unstructured and uncontrolled nature, can be enhanced and rendered more useful by robust terminological data. The project finished a few months ago – and a final report is eagerly anticipated, particularly as my formative research group submitted a proposal to JISC but lost out to EnTag! C'est la vie!