Thursday, 30 April 2009

WolframAlpha and destructive hype

If you have been plugged into the search engine or technology news feeds over recent months you may have encountered the excitement surrounding WolframAlpha. WolframAlpha is a Web search tool to be launched in May which – apart from having a good name – is anticipated to be the next Google exterminator. Although being touted as a destroyer of Google, the technology commentators indicate that WolframAlpha will inhabit an entirely different intellectual space on the Web.

WolframAlpha is described by its creators as a "computational knowledge engine" which, instead of retrieving resources using conventional automatic indexing methods, dynamically computes the answers to a wide variety of questions. The way in which it does this remains a mystery, but we do know that it models particular areas of knowledge. It then combines this with a vast repository of curated data harvested from disparate data sources and some ingenious natural language processing algorithms to represent knowledge. These knowledge representations can then be queried to answer real questions. Stills sounds like an enigma; but it must work on some level given the hype around it. Mustn't it?!

The brainchild of Dr. Stephen Wolfram (purveyor of computer algebra), WolframAlpha has had information and computer scientists and technology commentators salivating for months. The trouble is that while the incessant hype continues, an increasing number of people (me, but some commentators) are growing increasingly cynical of its true capabilities; we want to see a demo, or some kind of prototype. Mindful that cynicism could be spreading, Wolfram unveiled his creation yesterday for the first time at the Harvard University Berkman Center for Internet & Society (via a sold-out Web cast - clip from YouTube below). This demonstration appears to have further stimulated the hype (judging by some headlines), but has simultaneously added to the increasing cynicism. Hype and 'vapourware' exasperates people. And this is where the hype could actually be death of WolframAlpha, rather than Google.

In reality, it certainly sounds like WolframAlpha is not out to compete with Google; but it doesn't matter, this is how it is being described in the media and WolframAlpha hasn't tried to dispel the myth. In his blog, Wolfram describes WolframAlpha as "a new paradigm for using computing and the Web". It immediately provides people with a Google yardstick and false expectations; most new users will not understand that WolframAlpha is an entirely different beast. But more importantly, it's setting WolframAlpha up for an almighty fall.

Remember Cuil? People also thought Cuil was going to change the face of searching but it failed. It was hotly anticipated and was hyped, arguably more, than WolframAlpha. This hype did it no favours when it crashed on its launch day. It's only been 9 months since Cuil was officially launched, yet we never hear about it, nor do any of us use it. In part, this is because its indexes are so poor. My LJMU profile page was updated on 08 November 2008, almost 6 months ago; yet, Cuil still returns this page as it was on 07 November 2008 as a result. This is extremely feeble when you consider that Microsoft Live Search refreshes its indexes every 20 days.

Along with the hype, this was Cuil's 'blind spot'. A blind spot is normally tolerated in the early days of an innovative Web tool, but inflated expectations breeds intolerance. WolframAlpha is bound to have its own blind spot; what will it be and will users be tolerant until it's fixed? Probably not. They therefore have to get it right on launch day.

The moral of this tale is simple. Hyperbole must end. It's destructive and in the long run it does nobody any favours.

Tuesday, 7 April 2009

Web 2.0? Show me the money!

Just a quick post... Today the Guardian blog reports on the financial woes of YouTube. I don't suppose we should be particularly surprised to learn that according to some news sources YouTube is due to drop $470 million this year. When this figure is compared to the $1.65 billion pricetag Google paid a couple of years ago we can appreciate the magnitude of their YouTube predicament. The majority of this loss is attributable to the failure of advertising to bring home the bacon; a recurring issue on this blog. But huge running costs, copyright and royalty issues have played their part too. Google is reportedly interested in purchasing Twitter, but surely their failure to monetise YouTube - a service arguably more monetiseable (?) than Twitter - should have the alarm bells ringing at Google HQ?

I find the current crossroads for many of these services utterly fascinating. I don't have any solutions for any of these ventures, other than to make sure you have a business model before starting any business. Would RBS give me a business loan without a business plan and a robust revenue model? Probably not. But then they are not giving loans out these days anyway...

Wednesday, 25 March 2009

ASK conundrum revisited...again!

I posted a blog about Google's eye tracking research last month. I'm loathed to discuss Google again lest the ISG blog becomes known as the unofficial Google blog; however, the latest post on the Official Google Blog is worthy of some comment...

You might recall another post I made regarding search engine research and development, particularly in the area of information retrieval (IR) aids for users. In this posting I summarised Belkin's research and theories regarding the Anomalous State of Knowledge (ASK). Most of this and subsequent research has sought to introduce IR aids for the user so that they can better solve their ASK conundrum. This assistance varies but often takes the form of query expansion (in its various permutations), browsable subject trees to stimulate query formulation, relevance feedback, and so forth. Providing such tools in systems based on automatic indexing is difficult, but we noted that some search engines have introduced some effective retrieval aids, all designed to alleviate the ASK problem. For example, Yahoo! provides its search assist tool, Clusty provides related concept clusters, and Ask provides other similar tools. Their accuracy in IR varies widely, but overall they prove useful to user. Unfortunately, we also noted that Google provides few user aids comparable to those above, arguably relying more on its PageRank algorithm. Not any longer...

Today Google launched some interface functionality not dissimilar to Yahoo! search assist and Clusty. Their assistance provides some suggested related searches and some extra result summary text for particular results. Receiving this assistance depends on the nature of your query, so have a look at this canned search: 'communism in Russia'. This isn't bad and is better than nothing; but does it really measure up to the aids provided by competing search engines? Compare the results for these canned searches and the IR aids provided for the user by the systems we've discussed already:
Google's attempts appear quite pedestrian by comparison. Yahoo! and Clusty, for example, make their aids readily available so that the user can affect changes in their information seeking behaviour, but Google's tools are far less visible, less detailed, and offer far less functionality. Since a lot of research indicates that many users will not scroll below the 'golden triangle' (i.e. to the bottom of the first result set), it is entirely feasible to think that these 'related search' aids will go unnoticed by the disoriented information seeker.

It is good to see Google deploying user query aids and reacting to developments in other IR systems, but it appears that it will be some time before Google can be said to alleviate users' Anomalous State of Knowledge.

Tuesday, 17 March 2009

Who's going to teach our "stuff"? and who's going to learn it?

Is this the right forum for this? We all know of the imminent and unwelcome restructuring facing us. Where does the future lie for this discipline, or can we even define what our discipline is? We're constantly reminded now that our HE degrees are products, our students are customers, and so what are we, retailers? Compared to many other "products" in this HE marketplace our products are relatively unpopular despite the fact that there are fewer universities providing what we do compared to a decade ago.
I constantly struggle to explain what it is we do and can therefore understand why students have difficulty in placing it in context. Perhaps a business school isn't the right environment but then neither is a computing department, or it doesn't seem to be, and we don't fit into education or anywhere else.
What is the long term future for this discipline, whatever it is?

It's St. Patrick's day (not St. Paddy's, or St. Pat's or Paddy's day) so I'm off for Guinness in the local.

Thursday, 19 February 2009

Text is the new GUI?

We've got a software student ( from another University) working on his final year project at my day job. He is busy adding a 'speech' interface to our Laboratory tracking system. The idea is that while the scientist has their hands in the fume cupboard they don't want to be messing with a mouse or a keyboard why not engage with the computer by voice. This is after all how they did it in the old future in the movies.

Alas the student has got a bit sidetracked into the excitement of speech recognition and synthesis, in an attempt to get him to some sort of conclusion of his project I have suggested that most of the academic benefit of the project could be got by just having a text input and output (although useless in the fume cupboard). Once you go there it starts you thinking about how we interact with systems by voice. We are of course now used to listening to SatNav and some folk order their phones to phone the wife.

While pondering this and Georges earlier post of Twitter Library fees, I was thinking about an article about how fans have put together Twitter accounts of their favourite T.V. characters so that we can see when they are having a sandwich during the week. It made me wonder whether we might soon be engaging with various systems through the power of text rather than super 3D graphical interfaces.

This will be a shame as much of the design thought in web design and business application design has been about mice and windows more or less. In Liverpool this has led to some success for the hybrid 'programmer/graphic designer', perhaps if we are going to deal in a flow of text it will be the Hybrid 'programmer/DJ' or at least 'programmer/Linguist' who lead the way.

We surround ourselves in an increasing sense of a flow of consciousness through Twitter/Facebook etc. Surely this is going to include a Twitter from machines.

"Your fridge is enjoying a quiet day."
"Your car is worrying that it's service is due this week."
"Your door notes that fido is standing at it and wants to go out."

This will lead us to a wish to push application outputs into twitter like streams for our apps to respond to our own twitters. Text (or speech) may be the new GUI. Of course we will have to know a lot more about parsing and extracting meaning and identity from these streams of conciousness.

Thursday, 12 February 2009

FOAF and political social graphs

While catching up on some blogs I follow, I noticed that the Semantic Web-ite Ivan Herman posted comments regarding the US Congress SpaceBook – a US political answer to Facebook. He, in turn, was commenting on a blog made by the ProgrammableWeb – the website dedicated to keeping us informed of the latest web services, mashups, and Web 2.0 APIs.

From a mashup perspective, SpaceBook is pretty incredible, incorporating (so far) 11 different Web APIs. However, for me SpaceBook is interesting because it makes use of semantic data provided via FOAF and the microformat, XFN. To do this SpaceBook makes good use of the Google Social Graph API, which aims to harness such data to generate social graphs. The Social Graph API has been available for almost a year but has had quite a low profile until now. Says the API website:
"Google Search helps make this information more accessible and useful. If you take away the documents, you're left with the connections between people. Information about the public connections between people is really useful -- as a user, you might want to see who else you're connected to, and as a developer of social applications, you can provide better features for your users if you know who their public friends are. There hasn't been a good way to access this information. The Social Graph API now makes information about the public connections between people on the Web, expressed by XFN and FOAF markup and other publicly declared connections, easily available and useful for developers."
Bravo! This creates some neat connections. Unfortunately – and as Ivan Herman regrettably notes - the generated FOAF data is inserted into Hilary Clinton’s page as a page comment, rather than as a separate .rdf file or as RDFa. The FOAF file is also a little limited, but it does include links to her Twitter account. More puzzling for me though is why the embedded XHTML metadata does not use Qualified Dublin Core! Let's crank up the interoperability, please!

Friday, 6 February 2009

Information seeking behaviour at Google: eye-tracking research

Anne Aula and Kerry Rodden have just published a posting on the Official Google Blog summarising some eye-tracking research they have been conducting on Google's 'Universal Search'. Both are active in information seeking behaviour and human-computer interaction research at Google and are well published within the related literature (e.g. JASIST, IPM, SIGIR, CHI, etc.).

The motivation behind their research was to evaluate the effect incorporation of thumbnail images and video within a research set has on user information seeking behaviour. Previous information retrieval eye-tracking research indicates that users scan results in order, scanning down their results until they reach a (potentially) relevant result, or until they decide to refine their search query or abandon the search. Aula and Rodden were concerned that the inclusion of thumbnail images might distract the "well-established order of result evaluation". Some comparative evaluation was therefore order of the day.
"We ran a series of eye-tracking studies where we compared how users scan the search results pages with and without thumbnail images. Our studies showed that the thumbnails did not strongly affect the order of scanning the results and seemed to make it easier for the participants to find the result they wanted."
A good finding for Google, of course; but most astonishing is the eye-tracking data. The speed with which users scanned result sets and the number of points on the interface they scanned was incredible. View the 'real time' clip below. A dot increasing in size denotes the length of time a user spent pausing at that specific point in the interface or result set. Some other interesting discoveries were made – the full posting is essential reading.

Friday, 30 January 2009

Wikipedia: the new Knol?

Like many people I use Wikipedia quite regularly to check random facts. The strange aspect of this behaviour is that once I find the relevant fact, I have to immediately verify its provenance by conducting subsequent searches in order to find corroborative sources. It makes one wonder why one would use it in the first place.

Wikipedia continues to be plagued by a series of high profile malicious edits. Unfortunately, many of these edits aren't necessarily malicious. They are just wrong or inaccurate. There are probably hundreds of thousands of inaccurate Wikipedia articles, perhaps just as many hosting malicious edits; but it takes high profile gaffs to affect real change. On the day of Barack Obama's inauguration, Wikipedia reported the deaths of West Virginia's Robert Byrd and Edward Kennedy, who had collapsed during the inaugural lunch. Both reports were false.

This event appears to have compelled Jimmy Wales into being more proactive in improving the accuracy and reliability of Wikipedia. Under his proposals many future changes to articles would need to be approved by a group of vetted editors before being published. For me this news is interesting, particularly as it emerges barely two weeks after Google announced that the 100,000th Knol had been created on their 'authoritative and credible' answer to Wikipedia: Knol. Six or seven months ago the discussions focussed on how Knol was the new Wikipedia, now it appears as if Wikipedia might become the new Knol. How bizarre is that?!

Tightening the editing rules of Wikipedia has been on the agenda before and in 2007 this blog discussed how the German Wikipedia was conducting experiments which saw only trusted Wikipedians verifying changes to articles. So, will tightening the editing of Wikipedia make it the new Knol? The short answer is 'no'. Some existing Wikipedia editors can already exert authoritarian control over particular articles and can – in some cases – give the impression that they too have an axe to grind on particular topics. 'Wikipinochets' anyone? Moreover, Knol benefits from its "moderated collaboration" approach, with Knols being created by subject experts whose credentials have been verified. Wikipedia isn't going anywhere near this. Guardian columnist, Marcel Berlins, is probably right about Wikipedia when he states:
"I don't think there's a way of telling what proportion of Wikipedia entries are deficient, whether because of the writer's bias, mischief or lack of knowledge. It's clear that a significant number are questionable, sufficient to lead us to suspect all entries. But to do the right thing - vetting all contributors or contributions - would be impractical and hugely expensive. There is no easy solution. We many just have to accept that Wikipedia's undoubted usefulness comes at the price of occasional - perhaps frequent - inaccuracy. That is a sad conclusion to reach about an encyclopedia."
Oh well, back to verifying random facts found on Wikipedia.

Wednesday, 24 December 2008

SKOS-ifying Knowledge Organisation Systems: a continuing contradiction for the Semantic Web

A few days ago Ed Summers announced on his blog that he was shutting down lcsh.info. For those that don't know, lcsh.info was a Semantic Web demonstrator developed by Ed with the expressed purpose of illustrating how the Library of Congress Subject Headings (LCSH) could be represented and its structure harnessed using Simple Knowledge Organisation Systems (SKOS). In particular, Ed was keen to explore issues pertaining to Linked Data and the representation of concepts using URIs. He even hoped that the URIs used would be Cool URIs, linking eventually to a bona fide LCSH service were one ever to be released. Sadly, it was not to be... The reasons remain unclear but were presumably related to IPR. As the lcsh.info blog entry notes, Ed was compelled to remove it by the Library of Congress itself. The fact that he was the LC's resident Semantic Web buff probably didn't help matters, I'm sure.

SKOS falls within my area of interest and is an initiative of the Semantic Web Deployment Working Group. In brief, SKOS is an application of RDF and RDFS and is a series of evolving specifications and standards used to support the use of knowledge organisation systems (KOS) (e.g. information retrieval thesauri, classification schemes, subject heading systems, taxonomies or any other controlled vocabulary) within the framework of the Semantic Web. The Semantic Web is many things of course; but it is predicated upon the assumption that there exists communities of practices willing and able to create the necessary structured data (generally applications of RDF) to make it work. This might be metadata, or it might be an ontology, or it might be a KOS represented in SKOS. The resulting data can then be re-used, integrated, interconnected, queried and is open. When large communities of practice fail to contribute, the model breaks down.

There is a sense in which the Semantic Web has been designed to bring out the schizophrenic tendencies within some quarters of the LIS community. Whilst the majority of our community has embraced SKOS (and other related specifications), can appreciate the potential and actively contributes to the evolution of the standards, there is a small coterie that flirts with the technology whilst simultaneously shirking at the thought of exposing hitherto proprietary data. It's the 'lock down' versus 'openness' contradiction again.

In a previous research post I was involved with the High-Level Thesaurus (HILT) research project and continue my involvement in an consultative capacity. HILT continues to research and develop a terminology web service providing M2M access to a plethora of terminological data, including terminology mappings. Such terminological data can be incorporated into local systems to improve local searching functionality. Improvements might include, say, implementing a dynamic hierarchical subject browsing tree, or incorporating interactive query expansion techniques as part of the search interface, for example. An important - and the original motivation behind HILT - is to develop a 'terminology mapping server' capable of ameliorating the "limited terminological interoperability afforded between the federation of repositories, digital libraries and information services comprising the UK Joint Information Systems Committee (JISC) Information Environment" (Macgregor et al., 2007), thus enabling accurate federated subject-based information retrieval. This is a blog so detail will be avoided for now; but, in essence, HILT is an attempt to provide a terminology server in a mash-up context using open standards. To make the terminological data as usable as possible and to expose it to the Semantic Web, the data is modelled using SKOS.

But what happens to HILT when/if it becomes an operational service? Will its terminological innards be ripped out by the custodians of terminologies because they no longer want their data exposed, or will the ethos of the model be undermined as service administrators permit only HE institutions or charitable organisations from accessing the data? This isn't a concern for HILT yet; but it is one I anticipated several years ago. And the sad experience of lcsh.info illustrates that it's a very real concern.

Digital libraries, repositories and other information services have to decide where they want to be. This is a crossroads within a much bigger picture. Do they want their much needed data put to a good use on the Web, as some are doing (e.g. AGROVOC, GEMET, UKAT)? Or do they want alternative approaches to supplant them entirely (i.e. LCSH)? What's it gonna be, punks???

Monday, 8 December 2008

Wikipedia censorship: allusions to 'Smell the Glove'?

Another Wikipedia controversy rages, this time over censorship. Over the past two days, the Internet Watch Foundation informed some ISPs that an article pertaining to an album by the 'classic' German heavy metal band, Scorpions, may be illegal. Leaving aside the fact that Scorpions is one of many groups to have similar imagery on their record sleeves (the eponymous 1969 debut album by Blind Faith, Eric Clapton's supergroup, being another obvious example), am I the only person to notice the similarities with fictional rockumentary, This Is Spinal Tap?

Like metal, censorship is a heavy topic; but I thought this tenuous linkage with This Is Spinal Tap might be a welcome distraction from the usual blog postings, which are necessarily academic. Those of you familiar with said film might recall the controversy surrounding the proposed (tasteless) art work for Spinal Tap's new album (Smell the Glove), which in the end gets mothballed owing to its indecent nature. Getting into trouble over sleeve art is part and parcel of being in a heavy metal band it would seem! Enjoy the winter break, people!

Friday, 5 December 2008

Some general musings on tag clouds, resource discovery and pointless widgets...

The efficacy of collaborative tagging in information retrieval and resource discovery has undergone some discussion on this blog in the past. Despite emerging a good couple of years ago – and like many other Web 2.0 developments – collaborative tagging remains a topic of uncertainty; an area lacking sufficient evaluation and research. A creature of collaborative tagging which has similarly evaded adequate evaluation is the (seemingly ubiquitous!) 'tag cloud'. Invented by Flickr (Flickr tag cloud) and popularised by delicious (and aren't you glad they dropped the irritating full stops in their name and URL a few months ago?), tag clouds are everywhere; cluttering interfaces with their differently and irritatingly sized fonts.

Coincidentally, a series of tag cloud themed research papers were discussed at one of our recent ISG research group meetings. One of the papers under discussion (Sinclair & Cardew-Hall, 2008) conducted an experimental study comparing the usage and effectiveness of tag clouds with traditional search interface approaches to information retrieval. Their work is welcomed since it constitutes one of the few robust evaluations of tag clouds since they emerged several years ago.

One would hate to consider tag clouds as completely useless – and I have to admit to harbouring this thought. Fortunately, Sinclair and Cardew-Hall found tag clouds to be not entirely without merit. Whilst they are not conducive to precision retrieval and often conceal relevant resources, the authors found that users reported them useful for broad browsing and/or non-specific resource discovery. They were also found to be useful in characterising the subject nature of databases to be searched, thus aiding the information seeking process. The utility of tag clouds therefore remains confined to the search behaviours of inexperienced searchers and – as the authors conclude - cannot displace traditional search facilities or taxonomic browsing structures. As always, further research is required...

The only thing saving tag clouds from being completely useless is that they can occasionally assist you in finding something useful, perhaps serendipitously. What would be the point in having a tag cloud that didn't help you retrieve any information at all? Answer: There wouldn't be any point; but this doesn't stop some people. Recently we have witnessed the emergence of 'tag cloud generation' tools. Such tools generate tag clouds for Web pages, or text entered by the user. Wordle is one such example. They look nice and create interesting visualisations, but don't seem to do anything other than take a paragraph of text and increase the size of words based on frequency. (See the screen shot of a Wordle tag cloud for my home page research interests.)


OCLC have developed their very own tag cloud generator. Clearly, this widget has been created while developing their suite of nifty services, such as WorldCat, DeweyBrowser, FictionFinder, etc., so we must hold fire on the criticism. But unlike Wordle, this is something OCLC could make useful. For example, if I generate a tag cloud via this service, I expect to be able to click on a tag and immediately initiate a search on WorldCat, or a variety of OCLC services … or the Web generally! In line with good information retrieval practice, I also expect stopwords to be removed. In my example some of the largest tags are nonsense, such as "etc", "specifically", "use", etc. But I guess this is also a fundamental problem with tagging generally...

OCLC are also in a unique position in that they have access to numerous terminologies. This obviously cracks open the potential for cross-referencing tags with their terminological datasets so that only genuine controlled subject terms feature in the tag cloud, or productive linkages can be established between tags and controlled terms. This idea is almost as old as tagging itself but, again, has taken until recently to be investigated properly. Exploring the connections between tags and controlled vocabularies is something the EnTag project is exploring, a partner in which is OCLC. In particular, EnTag (Enhanced Tagging for Discovery) is exploring whether tag data, normally typified by its unstructured and uncontrolled nature, can be enhanced and rendered more useful by robust terminological data. The project finished a few months ago – and a final report is eagerly anticipated, particularly as my formative research group submitted a proposal to JISC but lost out to EnTag! C'est la vie!

Friday, 28 November 2008

Catching up with the old future of databases

We have been discussing in the group what we should be teaching on our Business Information Systems undergraduate course. Which is having a bit of a revamp. One area of discussion is about the areas of 'databases' which we teach mostly in the students second year and 'object oriented analysis and design' (but mostly analysis) which we teach in their final year.

When I used to teach database sections 10 years ago to business students we used to teach a history of:-
  1. file access
  2. hierarchical databases
  3. network databases
  4. relational databases
  5. object oriented databases.
Of course Object Oriented Databases hadn't happened in a big way back then. The surprising thing is that they haven't happened in a big way even now, they have spent 10 years being the next big thing. Meanwhile their close relatives UML based analysis and object oriented design and development have swept in from all directions. Strange that my notes from 10 years back now look like 'Space 1999', In their predictive powers. In my day job at Village we had a book on the subject 10 years back but a quick check of the book shelf shows we have long since recycled it.

Apparently believers and developers of Object Oriented Databases are just bemused by why everybody hasn't followed them into the promised land, particularly as much time is spent on 'Object Relational Mapping' technologies.

All round software development and architecture thinker and general purpose bearded Guru Martin Fowler, believes that the issue isn't to do with the general capabilities of the Object Oriented Databases but rather the fact that much integration in corporations occurs in the data layer not the business layer hence systems are dependent on standardised SQL approaches to integration. He suggests in his blog post on the subject that this shared database integration requirement has been holding back the march to the future of Object Oriented Databases. Creating extra inertia. Indeed, a confession, in my day job despite being Object Oriented N-Tier architecture developers by trade and conviction, when it came to tying our own timesheet system to our task management system we used database level triggers. It's a bit like the fact that there are better ways to do typing that using qwerty but we've all learnt to live with qwerty.

However with the movement towards using Web Services and SOA type architectures in effect making XML the linguq franca rather than SQL, Martin Fowler suggests that the field might start to loosen up. Although I wonder whether reporting is another issue. We produce some reports (in Crystal Reports and the equivalent) straight from our business objects, but other management reports really need to be produced straight off the (SQL) database. On occasions this is seperate from the main thrust of the application using a different technology stack.

Really the technology of the day in software development is Object Relational Mapping tools, ORMS. These try and hold the Object Oriented businesss layer to the Data Entity Oriented database layer. Such connections are relatively straightfoward in an unsophisticated design. My final year students currently angsting over a UML assignment will find that their Class Diagram is much the same as their Entity Relationship Diagram. But as you move deeper into doing things the Object way the two diverge. My Village colleague Ian Bufton and I have been discussing this in terms of lining up the two layers using either tools or code generation you can see some of his initial ponderings on his blog.

Luckily these types of contemplation are outside the scope of the things that our Information Systems students at the Business School, no doubt they have to worry about how to teach it at the Computer Science department.

Wednesday, 26 November 2008

'Gluing' searches with Yahoo!: part three in the search engine trilogy

There is plenty to comment on in the world of search engines at the moment. This post signifies the last in a series of discussions regarding search engines developments (or lack of!). (Part I; Part II)

As Google SearchWiki was unveiled, Yahoo! announced the wider release of Yahoo! Glue. Originally developed and tested at Yahoo! India, Yahoo! Glue now has a wider release – although it remains (perpetually?) in 'beta'. Glue is an attempt to aggregate disparate forms of information on a single results page in response to a single query. I suppose Glue is a functional demonstration of the ultimate mashup information retrieval tool. Glue assembles heterogeneous information from all over the Web, including text, news feeds, images, video, audio, etc. Search for the Beatles and you will get a results page listing Wikipedia definitions, LastFM tracks for your listening pleasure, news feeds, YouTube videos, etc.

In an ironic twist, Glue appears to be defying the dynamic ethos of Web 2.0. Glue searches are not created on the fly and only a limited number Glue searches are available at the moment (for example, no Liverpool!). Glue has the 'beta forever' mantra as the 'Get out of jail free card', of course. Still, Yahoo! informs us that:
"These pages are built using an algorithm that automatically places the most relevant modules on a page, giving you a visually rich, diverse page all about the topic in which you're interested."
Glue is also an example of Yahoo! exploring the social web in retrieval, harnessing as it does users' opinions on the accuracy of this algorithm (e.g. irrelevant or poorly ranked results can be 'flagged' as inappropriate or irrelevant).

Glue is - and will be - for the leisure user; the person falling into the 'popular search' category in search engines. These are the users submitting the simplest queries. The teenagers searching for 'Britney Spears', and the adults searching for 'Barack Obama' or 'Strictly Come Dancing'. The serious user (e.g. student, academic, knowledge worker, etc.) need not apply. I also have reservations over whether the summarisation of results is appropriate, and whether Glue can actually assemble disparate resources that are all relevant to a query. Check out these canned searches for Stephen Stills and Glasgow. In what way is Amsterdam a relation to Glasgow? And why the spurious news stories for Stephen Stills? Examining the text it is clear how it has been retrieved; but why when similar issues do not affect Yahoo! Search?

Like Google with SearchWiki, Yahoo! emphasise that Glue is not a replacements for Yahoo! Search; rather it's a "standalone experience":
"… Yahoo! Glue(TM) beta is not to replace the Yahoo! Search experience [...] We're always challenging ourselves to explore innovative new ways to deliver great experiences. Glue is one of those experiments, with a goal of giving users one more visual way to browse and discover new things from across the Web. We'll be working to expand the number of Glue pages, improve the experience and incorporate your feedback into future versions."
Very good. But making this dynamic and scalable should be atop the Glue 'to do' list. No Liverpool!

Monday, 24 November 2008

Wikifying search

This blog follows a series of other blogs pontificating about the efficacy of search engines in information retrieval. Over the weekend Google announced the release of Google SearchWiki. Google SearchWiki essentially allows users to customise searches by re-ranking, deleting, adding, and commenting on their results. This is personalised searching (see video below). As the Official Google Blog notes:
"With just a single click you can move the results you like to the top or add a new site. You can also write notes attached to a particular site and remove results that you don't feel belong."
The advantages of this are a little unclear at first; however, things become clearer when we learn that such changes can only be affected if you have an iGoogle account. Google have – quite understandably – been very specific about this aspect of SearchWiki. Search is their bread and butter; messing with the formula would be like dancing with the devil!

Google SearchWiki doesn't do anything further to address our Anomalous State of Knowledge (ASK), nor can I see myself using it, but it is an indication that Google is interested in better exploring the potential of social data to improve relevance feedback. Google will, of course, harvest vast amounts of data pertaining to users' information seeking behaviour which can then be channelled into improving their bread and butter. (And from my perspective, I would be interested to know how they analyse such data and affect changes in their PageRank algorithm). Their move also resonates with an increasing trend to support users in their Personal Information Management (PIM); to assist users in re-finding information they have previously located, or those frequently conducting the same searches over and over. It particularly reminds me of research undertaken by Bruce et al. (2004). For example, users increasingly chose not to bookmark a useful or interesting web page, but simply find it again – because they know they can. If you continually encounter information that is irrelevant to your area, re-rank it accordingly - so the SearchWiki ethos goes...

Perusing recent blogs it is clear that some consider this development to have business motivations. Technology guru and Wired magazine founder, John Battelle, thinks SearchWiki is an attempt to attract more users of iGoogle (which at the moment is small), whilst simultaneously rendering iGoogle the centre of users’ personal web universe. To my mind Google is always about business. PageRank is a great free-text searching tool, thus permitting huge market penetration. SearchWiki is simply another business tool which happens to offer some (vaguely?) useful functionality.

Tuesday, 7 October 2008

Search engines: solving the 'Anomalous State of Knowledge'

Information retrieval (IR) remains one of the most active areas of research within the information, computing and library science communities. It also remains one of the sexiest. The growth in information retrieval sex appeal has a clear correlation with the growth of the Web and the need for improvements in retrieval systems based on automatic indexing. No doubt the flurry of big name academics and Silicon Valley employees attending conferences such as SIGIR also adds glamour. Nevertheless, the allure of IR research has precipitated some of the best innovations in IR ever, as well as creating some of the most important search engines and business brands. Of course, asked to pick from a list their favourite search engine or brand, most would probably select Google.

The habitual use of Google by students (and by real people generally!) was discussed in a previous post and needn't be revisited here. Nevertheless, one of the most distressing aspects of Google (for me, at least!) is a recent malaise in its commitment to search. There have been some impressive innovations in a variety of search engines in a variety of areas. For example, Yahoo! is to better harness metadata and Semantic Web data on the Web. More interestingly though, some recent and impressive innovations in solving the 'ASK conundrum' is visible in a variety of search engines, but not in Google. Although Google always tell us that search is its bread and butter, is it spreading itself a little too thinly? Or - with a brand loyalty second to none and the robust PageRank algorithm deployed to good effect – is Google resting on its laurels?

In 1982 a young Nicholas J. Belkin spearheaded a series of seminal papers documenting various models of users' information needs in IR. These papers remain relevant today and are frequently cited. One of Belkin et al.'s central suppositions is that the user suffers from the so-called Anomalous State of Knowledge, which can be conveniently acronymized to 'ASK'. Their supposition can be summarised by the following quote from their JDoc paper:
"[P]eople who use IR systems do so because they have recognised an anomaly in their state of knowledge on some topic, but they are unable to specify precisely what is necessary to resolve that anomaly. ... Thus, we presume that it is unrealistic (in general) to ask the user of an IR system to say exactly what it is that she/he needs to know, since it is just the lack of that knowledge which has brought her/him to the system in the first place".
This astute deduction ushered in a branch of IR research that sought to improve retrieval by resolving the Anomalous State of Knowledge (e.g. providing the user with assistance in the query formulation process, helping users ‘fill in the blanks’ to improve recall (e.g. query expansion), etc.).

Last winter Yahoo! unveiled its 'Search Assist' facility (see screenshot above - search for 'united nations'), which provides a real time query formulation assistance to the user. Providing these facilities in systems based on metadata has always been possible owing the use of controlled vocabularies for indexing, the use of name authority files, and even content standards such as AACR2; but providing a similar level functionality with unstructured information is difficult – yet Yahoo! provide something ... and it can be useful and can actually help resolve the ASK conundrum!


Similarly, meta-search engine Clusty has provided its 'clustering' techniques for quite some time. These clusters group related concepts and are designed to aid in query formulation, but also to provide some level of relevance feedback to users (see screenshot above - search for 'George Macgregor'). Of course, these clusters can be a bit hit or miss but, again, they can improve retrieval and aid the user in query formulation. Similar developments can also be found in Ask. View this canned search, for example. What help does Google provide?

The bottom line is that some search engines are innovating endlessly and putting the fruits of a sexy research area to good use. These search engines are actually moving search forward. Can the same still be said of Google?