Friday, 12 June 2009

Quasi-facetted retrieval of images using emotions?

As part of my literature catch up I found an extremely interesting paper in JASIST by S. Schmidt and Wolfgang G. Stock entitled, 'Collective indexing of emotions in images : a study in emotional information retrieval'. The motivation behind the research is simple: images tend to elicit emotional responses in people. Is it therefore possible to capture these emotional responses and use them in image retrieval?

An interesting research question indeed, and Schmidt and Stock's study found that 'yes', it is possible to capture these emotional responses and use them. In brief, their research asked circa 800 users to tag a variety of public images from Flickr using their scroll-bar tagging system. This scroll-bar tagging system allowed users to tag images according to a series of specially selected emotional responses and to indicate the intensity of these emotions. Schmidt and Stock found that users tended to have favourite emotions and this can obviously differ between users; however, for a large proportion of images the consistency of emotion tagging is very high (i.e. a large proportion of users frequently experience the same emotional response to an image). It's a complex area of study and their paper is recommended reading precisely for this reason (capturing emotions anyone?!), but their conclusions suggest that:
"…it seems possible to apply collective image emotion tagging to image information systems and to present a new search option for basic emotions."
To what extent does the image above (by D Sharon Pruitt) make you feel happiness, anger, sadness, disgust or fear? It is early days, but the future application of such tools could find a place within the growing suite of image filters that many search engines have recently unveiled. For example, yesterday Keith Trickey was commenting on the fact that the image filters in Bing are better than Google or Yahoo!. True. There are more filters, and they seem to work better. In fact, they provide a species of quasi-taxonomical facets: (by) size, layout, color, style and people. It's hardly Ranganathan's PMEST, but – keeping in mind that no human intervention is required - it's a useful quasi-facet way of retrieving or filtering images, albeit flat.

An emotional facet, based on Schmidt and Stock's research, could easily be added to systems like Bing. In the medium term it is Yahoo! that will be more in a position to harness the potential of emotional tagging. They own Flickr and have recently incorporated the searching and filtering of Flickr images within Yahoo! Image Search. As Yahoo! are keen for us to use Image Search to find CC images for PowerPoint presentations, or to illustrate a blog, being able to filter by emotions would be a useful addition to the filtering arsenal.

Thursday, 11 June 2009

Bada Bing!

So much has been happening in the world of search engines since spring this year. This much can be evidenced from the postings on this blog. All the (best) search engines have been active in improving user tools, features, extra search functionality, etc. and there is a real sense that some serious competition is happening at the moment. It's all exciting stuff…

Last week Microsoft officially released its new Bing search engine. I've been using it, and it has found things Google hasn't been able to. The critics have been extremely impressed by Bing too and some figures suggest that it is stealing market share and moving Yahoo! to the number 2 spot. What about number 1?

The trouble is that it doesn't matter how good your search engine is because it will always have difficulty interrupting users' habitual use of Google. Indeed, Google's own research has demonstrated that the mere presence of the Google logo atop a result set is a key determinant of whether a user is satisfied with their results or not. In effect, users can be shown results from Yahoo! but branded as Google, and vice versa, but will always choose the result with the Google branding. Thus, users are generally unable to tell whether there is any real difference in the results (i.e. their precision, relevance, etc.) and are actually more influenced by the brand and their past experience. It's depressing, but a reality for the likes of Microsoft, Yahoo!, Ask, etc.

Francis Muir has the 'Microsoft mantra'. He predicts that in the long run Microsoft is always going to dominate Google – and I am starting to agree with him. Microsoft sit back, wait for things to unfold, and then develop something better than its previously dominant competitors. True – they were caught on the back foot with Web searching, but Bing is as at least as good as Yahoo!, perhaps better, and it can only get better. Their contribution to cloud computing (SkyDrive) offers 25GB storage, integration with Office and email, etc. and is far better than anything else available. Google documents? Pah! Who are you going to share that with? And then you consider Microsoft's dominance in software, operating systems, programming frameworks, databases, etc. Integrating and interoperating with this stuff over the Web is a significant part of the Web's future. Google is unlikely to be part of this, and for once I'm pleased.

It is not Microsoft's intention to take on Google's dominance of the Web at the moment. But I reckon Bing is certainly part of the long term strategy. The Muir prophecy is one step closer methinks.

Cracking open metadata and cataloguing research with Resource Description & Access (RDA)

I have been taking the opportunity to catch up with some recently published literature over the past couple of weeks. While perusing the latest issue of the Bulletin of the American Society for Information Science and Technology (the magazine which complements JASIST), I read an interesting article by Shawne D. Miksa (associate professor at the College of Information, University of North Texas). Miksa's principal research interests reside in metadata, cataloguing and indexing. She has been active in disseminating about Resource Description & Access (RDA) and has a book in the pipeline designed to demystify it.

RDA has been in development for several years now, is the successor to AACR2 and provides rules and guidance on the cataloguing of information entities. I use the phrase 'information entities' since RDA departs significantly from AACR2. The foundations of AACR2 were created prior to the advent of the Web and this remains problematic given the digital and new media information environment in which we now exist. Of course, more recent editions of AACR2 have attempted to better accommodate these developments, but fire fighting was always order of the day. The now re-named Joint Steering Committee for the Development of RDA has known for quite some time that an entirely new approach was required – and a few years ago radical changes to AACR2 were announced. As my ex-colleague Gordon Dunsire describes in a recent D-Lib Magazine article:
"RDA: Resource Description and Access is in development as a new standard for resource description and access designed for the digital world. It is being built on the foundation established for the Anglo-American Cataloguing Rules (AACR). Although it is being developed for use primarily in libraries, it aims to attain an effective level of alignment with the metadata standards used in related communities such as archives, museums and publishers, and to provide a better fit with emerging database technologies."
The ins and outs of RDA is a bit much for this blog; suffice to say that RDA is ultimately designed to improve the resource discovery potential of digital libraries and other retrieval systems by utilising the FRBR conceptual entity-relationship model (see this entity-relationship diagram at the FRBR blog). FRBR provides a holistic approach to users' retrieval requirements by establishing the relationships between information entities and allowing users to traverse the hierarchical relationships therein. I am an advocate of FRBR and appreciate its retrieval potential. Indeed, I often direct postgraduate students to Fiction Finder, an OCLC Research prototype which demonstrates the FRBR Work-Set Algorithm.

Reading Miksa's article was interesting for two reasons. Firstly, RDA has fallen off of my radar recently. I used to be kept abreast of RDA development through the activities of my colleague Gordon, who also disseminates widely on RDA and feeds into the JSC's work. Miksa's article – which announces the official release of RDA in second half of 2009 – was almost like being in a time machine! RDA is here already! Wow! It only seems like last week when JSC started work on RDA (...but it was actually over 5 years ago…).

The development of RDA has been extremely controversial, and Miksa alludes to this in her article – metadata gurus clashing with traditional cataloguers clashing with LIS revolutionaries. It has been pretty ugly at times. But secondly – and perhaps more importantly – Miksa's article is a brilliant call to arms for more metadata research. Not only that, she notes areas where extensive research will be mandatory to bring truly FRBR-ised digital libraries to fruition. This includes consideration of how this impacts upon LIS education.

A new dawn? I think so… Can the non-believers grumble about that? Between the type of developments noted earlier and RDA, the future of information organisation is alive and kicking.

Thursday, 4 June 2009

Fight! Google Squared vs. WolframAlpha

By now we all realise that WolframAlpha is not intended to compete with Google's Universal Search; it's a 'computational knowledge engine' designed to serve up facts, data and scientific knowledge and is an entirely different beast. Nevertheless, Google is not a company to be outdone and has just announced the release of Google Squared which, if the technology press is to be believed, is Google's attempt to usurp WolframAlpha's grip on offering up facts, data and knowledge. Indeed, Google attempted to steal WolframAlpha's thunder by announcing that Google Squared was in development on the same day Stephen Wolfram was unveiling WolframAlpha for the first time a few weeks ago. Meow!

In the same way that WolframAlpha occupies a different intellectual space to most web search engines, Google Squared seems to be quite different to WolframAlpha. Says the Official Google Blog:
"Google Squared is an experimental search tool that collects facts from the web and presents them in an organized collection, similar to a spreadsheet. If you search for [roller coasters], Google Squared builds a square with rows for each of several specific roller coasters and columns for corresponding facts, such as image, height and maximum speed."
Google Squared appears to work best when the query submitted is conducive to comparing species of, say, snakes or country rock bands. With the former you retrieve a variety of snake types, images, description, as well as biological taxonomic classification data, and with the latter genre and date of band formation is retrieved (including Dillard & Clark and the Flying Burrito Brothers), in addition to images and descriptions. Many of the data values are incorrect, but Google has been quite forthright in stating that Google Squared to extremely experimental ("This technology is by no means perfect"; "Google Squared is an experimental search tool"). Of course, Google wants us to explore their canned searches, such as Rollercoasters or African countries, to best appreciate what is possible.

As we noted recently though, place names are good to test these systems and, like WolframAlpha, some bizarre results are retrieved. A search for Liverpool seems to only retrieve facts on assorted Liverpool F.C. players, and Glasgow retrieves persons associated with Glasgow and the death of Glasgow Central train station in 1989(!) I had hoped Google Squared's comparative power might have pulled together facts and statistics on Glasgow (UK) with the ten or so places named Glasgow in the USA and Canada. A similar result would have been expected for Liverpool or Manchester (which has far, far more), but alas. This is a particular shame owing to the fact that much of this data is available on Wikipedia in a relatively structured format, with disambiguation pages to help.

Google Squared allows users to correct data, remove erroneous results or suggest better results. The effect of this is a dynamically evolving result set. A search for a popular topic an hour ago can yield an entirely different result an hour later. All of this will help Google Squared become more accurate and cleverer over time.

Although Google Squared and WolframAlpha are quite different, there are some similarities. For this reason it is possible to state that the current score is 1-0 to WolframAlpha.

Wednesday, 27 May 2009

Image searching with Creative Commons

Student information literacy skills have been discussed on the blog before. In short, they are woeful. One area where students tend to have little understanding is in the area of intellectual property rights (IPR). The situation might be looking better for digital music, but in my experience it remains poor for other digital artefacts, particularly images. 'Twas only a few weeks ago while was in a lab with some undergraduate students for a web technologies module when I discovered most of them were ripping images from the web for inclusion within their information gateways. While this can (in some circumstances) be tolerated within the confines of an educational institution, it remains copyright infringement owing to copying by 'reprographic means' - and this isn't behaviour we want to become habitual in our graduates. My brother (a graphic designer and new media guru) has spun me many a yarn about ex-colleagues who have been shown their P45 for engaging in IPR theft (e.g. reusing someone's basic design or photograph).

All of this is veering away from the original reason for this blog though, which is to draw attention to some new image searching functionality on Yahoo! Image Search. Following on nicely from the Search Options post, the Yahoo! Search Blog has just announced the inclusion of some extra search filters for image result sets. Not only is it better than Google (and more accurate?), but it also includes a useful Creative Commons (CC) filter. Using a similar interface to Yahoo! Search Assist, Yahoo! Image Search allows users to apply a CC checkbox to filter for images, with specific filters included for commercial reuse and/or remixing. This is particularly useful to embellish those PowerPoint presentations or to illustrate a blog, or for those undergraduate students building an information gateway, or to avoid getting a P45!

There appears to be a downside, unfortunately. When I saw the Yahoo! Search Blog announcement I thought (perhaps naively) that Yahoo! was starting to put into practice its commitment to metadata, Semantic Web specifications, and other structured data. Since I know my personal homepage is indexed by Yahoo! and uses XHTML+RDFa to notify intelligent agents that its page content falls under a Creative Commons Attribution 3.0 License, I thought I'd put an Image Search to the test. Providing the CC namespace is referenced, the XHTML+RDFa required is simple. For example:

<p>Content on <a href="http://www.staff.ljmu.ac.uk/bsngmacg/" property="cc:attributionName" rel="cc:attributionURL">George Macgregor</a>'s website is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by/3.0/">Creative Commons Attribution 3.0 License</a></p>

...and with specific CC reference to my foaf:depiction...

<img src="img/georgedepiction.jpg" alt="Image of George Macgregor" rel="license" href="http://creativecommons.org/licenses/by/3.0/" property="foaf:depiction" content="George Macgregor"/>

My filtered CC search was unsuccessful though. This disappointed me; but then I observed the following notice:
"Note: Only Flickr images are supported currently."
Flickr – which is a subsidiary of Yahoo! – has allowed users to conduct advanced searches of its publicly uploaded images for quite some time. This has included CC searching. And it would appear that Yahoo! has integrated Flickr searching functionality into Image Search, albeit with some nice tweaks. If I had read their blog in its entirety I would have realised this; I just clicked the link such was my excitement about Yahoo! Image Search!

It's useful to have this functionality within a conventional searching tool, but it is disappointing that Image Search isn't using cleverer means of doing it (e.g. RDFa) rather than relying on the preferences of Flickr users when they upload their images. Don't get me wrong, this is useful and most welcome, and it will save me time on occasion, but it would be exciting to crack CC image searching beyond the controlled Flickr environment. Hopefully the 'currently' in "Only Flickr images are supported currently" will mean that my expectations will be met soon…

Monday, 18 May 2009

Light relief: Celtic fringes erased by WolframAlpha?!

With WolframAlpha launched on Friday, I spent much of my weekend trying to get a 'computable request' to compute. Not until Monday morning did a request compute – but its performance has been getting better ever since so hopefully we will all have more time to experiment with it over coming days and weeks...

Like me, Gwenda Mynott has been testing WolframAlpha and has been searching for things that, a) you have a good knowledge of, and, b) a topic that WolframAlpha can easily compute. Places are good for this (e.g. countries, towns, cities, etc.), and Stephen Wolfram computes multiple locations to good effect in his demonstrations; however, Gwenda tried to 'compute' Wales and arrived at some bizarre results. Check them out. WolframAlpha doesn't retrieve data pertaining to the constituent nation of the United Kingdom of Great Britain and Northern Ireland (i.e. Wales as you or I would tend to know it!), but a small town in South Yorkshire by the name of Wales (?) The only other obvious option WolframAlpha provides is Wales (New York, USA), which is equally amiss.

Hmmmmm. If this is the result for Wales, what are the results for the rest of the UK? Well, that's equally controversial. England appears to be synonymous with the United Kingdom of Great Britain and Northern Ireland. Scotland is referred by WolframAlpha back to the Kingdom of Scotland, which ceased to exist after the Act of Union in 1707. Worse than that, Northern Ireland doesn't even exist! ("WolframAlpha isn't sure what to do with your input") Cornish nationalists will also be dismayed to learn that Cornwall (Canada) is the only one that counts.

Is this a systematic attempt to erase the history, culture and memory of the Celtic fringes?! Of course not. The results might be strange, but from a knowledge engine point of view – and ontologically speaking - Wales, Scotland and Northern Ireland are subsumed by the larger geographical and political entity of the UK, so it's understandable that WolframAlpha computes the answer in this way. Still, the England/UK synonymy is a bit odd and must have been encoded by someone somewhere sometime!

Experiment away, folks - and I would encourage everyone to post their most bizarre / illogical data results as comments to this blog. A prize will go to the most outlandish!

Friday, 15 May 2009

Some more 'Search Options'...

I promised not to blog about Google any time soon for fear the blog becomes known as the unofficial Google blog. After some consideration I thought, 'pish posh!' Anyway, the post has a wider remit than just Google...honest!

The absence of retrieval aids for Google users (oh no, not again - I hear you cry!) has been discussed at great length on this blog before. To appreciate the extent of this deficiency we need only peruse some innovative rival search engines such as Ask (recently re-branded back to Ask Jeeves), Yahoo!, or Clusty. Google has been making changes though and today the Official Google Blog announced some further enhancements to the universal Google search interface. Simply called, 'Search Options', these tools let you "slice and dice" results, apply rudimentary filters, and generate alternative views of results. Search Options does a little bit more to help the user in query formulation (the area where I think Google is weakest), but also offers some useful functionality once you have your results.



Check out our usual canned search for 'communism in Russia'; 'click' the 'Search options…' link in the top left had corner of the interface to reveal the Search Option tools.

Filters are available for videos, forums and reviews (the latter being fairly useful if you are shopping). Various publication time filters are also available. Nothing here is particularly mind blowing though.

Search Options gets a bit more interesting when the search display options are explored in a little detail. Firstly, it's possible to request details of related searches. These are displayed in a better page location than before and look similar to Yahoo! Search Assist. But it is now also possible to select the 'Wonder Wheel' which generates a visualisation of the related terms. I'm unsure how useful the Wonder Wheel really is, particularly as the true nature of the relationships between terms is impossible for Google to represent other than in syntactic terms; this is something the Semantic Web community is obviously trying to resolve.

Most interesting though is the 'Timeline' tool. This allows results to be displayed along, erm, yup, a timeline. The timeline is clickable allowing the user to drill down into particular temporal zones and to view resources relating to that zone. I use the word 'interesting' because although the timeline is probably quite useful for historical research, its moment of introduction is the most interesting part. Indeed, the timeline functionality looks in part like Google is bracing itself for the release of WolframAlpha, which is due any day now (or tonight?) – and I wouldn't be at all surprised if this announcement was an attempt to steal some of its thunder. This appears to have been combined with the demonstration of Google Squared at the Google Searchology conference a few days ago. No Google Squared prototypes appear to be available for us to experiment with, but TechCrunch got a sneaky peak at Searchology (view the YouTube video below). Google Squared is, in essence, Google's answer to WolframAlpha.

For me the most interesting news to emerge alongside Search Options is Google's desire to make greater use of RDFa. RDFa is probably a little pedestrian for me, but it's better than nothing – and at least there is a clear intention of using some Semantic Web specifications. It's just a shame Yahoo! announced something similar but more radical almost 18 months ago.

Friday, 1 May 2009

LCSH as Linked Data ... officially!

Yesterday was, in my estimation, pretty historic. The Library of Congress officially launched the LC Authorities and Vocabularies service. You might recall a previous post relating to lcsh.info in which I lamented the LC's decision to pull down a SKOS demonstrator of LCSH, explicitly designed to explore the possibilities of Linked Data and dereferenceable URIs. All the background is in the previous post; but the whole episode appears to have been a PR disaster for LC.

The great news is that the LC Authorities and Vocabularies service (let's call it LCAV henceforth, shall we?) officially re-launched lcsh.info in a bigger, better and much improved form. The service essentially enables both humans and machines to access a plethora of LC authority data. Like lcsh.info, the service employs Semantic Web approaches to exposing this data and implements approaches to Linked Data by exposing and linking data on the Web via dereferenceable URIs.

Five minutes exploring the website reveals that LCAV serves up the entire LCSH for free, with incredible search and browse functionality, leaving Connexion in the shade. The concept URIs point to detailed data modelled in SKOS as RDFa for human readability, but with links to SKOS as RDF/XML, N-Triples and (the less familiar?) JSON for machine processing. RDF graphs can even be visualised by clicking, well, the 'visualize' tab – incredible. Mappings to other vocabularies are also provided.
On top of all this, LCSH can be downloaded in its entirety as RDF/XML or N-Triples (SKOS)! LCAV also indicate that further authority data will be made available soon.

Make no bones about it, this is historic stuff, not only because the service is so good but because this terminological data is no longer locked down. I think it's important to stroke our imaginary beards over the significance of the LC's change of direction. Is this the beginning of the end for locked down terminological data?! Will they be like dominoes henceforth? A fiver says DDC does the same by the end of the year. Any takers???

Thursday, 30 April 2009

WolframAlpha and destructive hype

If you have been plugged into the search engine or technology news feeds over recent months you may have encountered the excitement surrounding WolframAlpha. WolframAlpha is a Web search tool to be launched in May which – apart from having a good name – is anticipated to be the next Google exterminator. Although being touted as a destroyer of Google, the technology commentators indicate that WolframAlpha will inhabit an entirely different intellectual space on the Web.

WolframAlpha is described by its creators as a "computational knowledge engine" which, instead of retrieving resources using conventional automatic indexing methods, dynamically computes the answers to a wide variety of questions. The way in which it does this remains a mystery, but we do know that it models particular areas of knowledge. It then combines this with a vast repository of curated data harvested from disparate data sources and some ingenious natural language processing algorithms to represent knowledge. These knowledge representations can then be queried to answer real questions. Stills sounds like an enigma; but it must work on some level given the hype around it. Mustn't it?!

The brainchild of Dr. Stephen Wolfram (purveyor of computer algebra), WolframAlpha has had information and computer scientists and technology commentators salivating for months. The trouble is that while the incessant hype continues, an increasing number of people (me, but some commentators) are growing increasingly cynical of its true capabilities; we want to see a demo, or some kind of prototype. Mindful that cynicism could be spreading, Wolfram unveiled his creation yesterday for the first time at the Harvard University Berkman Center for Internet & Society (via a sold-out Web cast - clip from YouTube below). This demonstration appears to have further stimulated the hype (judging by some headlines), but has simultaneously added to the increasing cynicism. Hype and 'vapourware' exasperates people. And this is where the hype could actually be death of WolframAlpha, rather than Google.

In reality, it certainly sounds like WolframAlpha is not out to compete with Google; but it doesn't matter, this is how it is being described in the media and WolframAlpha hasn't tried to dispel the myth. In his blog, Wolfram describes WolframAlpha as "a new paradigm for using computing and the Web". It immediately provides people with a Google yardstick and false expectations; most new users will not understand that WolframAlpha is an entirely different beast. But more importantly, it's setting WolframAlpha up for an almighty fall.

Remember Cuil? People also thought Cuil was going to change the face of searching but it failed. It was hotly anticipated and was hyped, arguably more, than WolframAlpha. This hype did it no favours when it crashed on its launch day. It's only been 9 months since Cuil was officially launched, yet we never hear about it, nor do any of us use it. In part, this is because its indexes are so poor. My LJMU profile page was updated on 08 November 2008, almost 6 months ago; yet, Cuil still returns this page as it was on 07 November 2008 as a result. This is extremely feeble when you consider that Microsoft Live Search refreshes its indexes every 20 days.

Along with the hype, this was Cuil's 'blind spot'. A blind spot is normally tolerated in the early days of an innovative Web tool, but inflated expectations breeds intolerance. WolframAlpha is bound to have its own blind spot; what will it be and will users be tolerant until it's fixed? Probably not. They therefore have to get it right on launch day.

The moral of this tale is simple. Hyperbole must end. It's destructive and in the long run it does nobody any favours.

Tuesday, 7 April 2009

Web 2.0? Show me the money!

Just a quick post... Today the Guardian blog reports on the financial woes of YouTube. I don't suppose we should be particularly surprised to learn that according to some news sources YouTube is due to drop $470 million this year. When this figure is compared to the $1.65 billion pricetag Google paid a couple of years ago we can appreciate the magnitude of their YouTube predicament. The majority of this loss is attributable to the failure of advertising to bring home the bacon; a recurring issue on this blog. But huge running costs, copyright and royalty issues have played their part too. Google is reportedly interested in purchasing Twitter, but surely their failure to monetise YouTube - a service arguably more monetiseable (?) than Twitter - should have the alarm bells ringing at Google HQ?

I find the current crossroads for many of these services utterly fascinating. I don't have any solutions for any of these ventures, other than to make sure you have a business model before starting any business. Would RBS give me a business loan without a business plan and a robust revenue model? Probably not. But then they are not giving loans out these days anyway...

Wednesday, 25 March 2009

ASK conundrum revisited...again!

I posted a blog about Google's eye tracking research last month. I'm loathed to discuss Google again lest the ISG blog becomes known as the unofficial Google blog; however, the latest post on the Official Google Blog is worthy of some comment...

You might recall another post I made regarding search engine research and development, particularly in the area of information retrieval (IR) aids for users. In this posting I summarised Belkin's research and theories regarding the Anomalous State of Knowledge (ASK). Most of this and subsequent research has sought to introduce IR aids for the user so that they can better solve their ASK conundrum. This assistance varies but often takes the form of query expansion (in its various permutations), browsable subject trees to stimulate query formulation, relevance feedback, and so forth. Providing such tools in systems based on automatic indexing is difficult, but we noted that some search engines have introduced some effective retrieval aids, all designed to alleviate the ASK problem. For example, Yahoo! provides its search assist tool, Clusty provides related concept clusters, and Ask provides other similar tools. Their accuracy in IR varies widely, but overall they prove useful to user. Unfortunately, we also noted that Google provides few user aids comparable to those above, arguably relying more on its PageRank algorithm. Not any longer...

Today Google launched some interface functionality not dissimilar to Yahoo! search assist and Clusty. Their assistance provides some suggested related searches and some extra result summary text for particular results. Receiving this assistance depends on the nature of your query, so have a look at this canned search: 'communism in Russia'. This isn't bad and is better than nothing; but does it really measure up to the aids provided by competing search engines? Compare the results for these canned searches and the IR aids provided for the user by the systems we've discussed already:
Google's attempts appear quite pedestrian by comparison. Yahoo! and Clusty, for example, make their aids readily available so that the user can affect changes in their information seeking behaviour, but Google's tools are far less visible, less detailed, and offer far less functionality. Since a lot of research indicates that many users will not scroll below the 'golden triangle' (i.e. to the bottom of the first result set), it is entirely feasible to think that these 'related search' aids will go unnoticed by the disoriented information seeker.

It is good to see Google deploying user query aids and reacting to developments in other IR systems, but it appears that it will be some time before Google can be said to alleviate users' Anomalous State of Knowledge.

Tuesday, 17 March 2009

Who's going to teach our "stuff"? and who's going to learn it?

Is this the right forum for this? We all know of the imminent and unwelcome restructuring facing us. Where does the future lie for this discipline, or can we even define what our discipline is? We're constantly reminded now that our HE degrees are products, our students are customers, and so what are we, retailers? Compared to many other "products" in this HE marketplace our products are relatively unpopular despite the fact that there are fewer universities providing what we do compared to a decade ago.
I constantly struggle to explain what it is we do and can therefore understand why students have difficulty in placing it in context. Perhaps a business school isn't the right environment but then neither is a computing department, or it doesn't seem to be, and we don't fit into education or anywhere else.
What is the long term future for this discipline, whatever it is?

It's St. Patrick's day (not St. Paddy's, or St. Pat's or Paddy's day) so I'm off for Guinness in the local.

Thursday, 19 February 2009

Text is the new GUI?

We've got a software student ( from another University) working on his final year project at my day job. He is busy adding a 'speech' interface to our Laboratory tracking system. The idea is that while the scientist has their hands in the fume cupboard they don't want to be messing with a mouse or a keyboard why not engage with the computer by voice. This is after all how they did it in the old future in the movies.

Alas the student has got a bit sidetracked into the excitement of speech recognition and synthesis, in an attempt to get him to some sort of conclusion of his project I have suggested that most of the academic benefit of the project could be got by just having a text input and output (although useless in the fume cupboard). Once you go there it starts you thinking about how we interact with systems by voice. We are of course now used to listening to SatNav and some folk order their phones to phone the wife.

While pondering this and Georges earlier post of Twitter Library fees, I was thinking about an article about how fans have put together Twitter accounts of their favourite T.V. characters so that we can see when they are having a sandwich during the week. It made me wonder whether we might soon be engaging with various systems through the power of text rather than super 3D graphical interfaces.

This will be a shame as much of the design thought in web design and business application design has been about mice and windows more or less. In Liverpool this has led to some success for the hybrid 'programmer/graphic designer', perhaps if we are going to deal in a flow of text it will be the Hybrid 'programmer/DJ' or at least 'programmer/Linguist' who lead the way.

We surround ourselves in an increasing sense of a flow of consciousness through Twitter/Facebook etc. Surely this is going to include a Twitter from machines.

"Your fridge is enjoying a quiet day."
"Your car is worrying that it's service is due this week."
"Your door notes that fido is standing at it and wants to go out."

This will lead us to a wish to push application outputs into twitter like streams for our apps to respond to our own twitters. Text (or speech) may be the new GUI. Of course we will have to know a lot more about parsing and extracting meaning and identity from these streams of conciousness.

Thursday, 12 February 2009

FOAF and political social graphs

While catching up on some blogs I follow, I noticed that the Semantic Web-ite Ivan Herman posted comments regarding the US Congress SpaceBook – a US political answer to Facebook. He, in turn, was commenting on a blog made by the ProgrammableWeb – the website dedicated to keeping us informed of the latest web services, mashups, and Web 2.0 APIs.

From a mashup perspective, SpaceBook is pretty incredible, incorporating (so far) 11 different Web APIs. However, for me SpaceBook is interesting because it makes use of semantic data provided via FOAF and the microformat, XFN. To do this SpaceBook makes good use of the Google Social Graph API, which aims to harness such data to generate social graphs. The Social Graph API has been available for almost a year but has had quite a low profile until now. Says the API website:
"Google Search helps make this information more accessible and useful. If you take away the documents, you're left with the connections between people. Information about the public connections between people is really useful -- as a user, you might want to see who else you're connected to, and as a developer of social applications, you can provide better features for your users if you know who their public friends are. There hasn't been a good way to access this information. The Social Graph API now makes information about the public connections between people on the Web, expressed by XFN and FOAF markup and other publicly declared connections, easily available and useful for developers."
Bravo! This creates some neat connections. Unfortunately – and as Ivan Herman regrettably notes - the generated FOAF data is inserted into Hilary Clinton’s page as a page comment, rather than as a separate .rdf file or as RDFa. The FOAF file is also a little limited, but it does include links to her Twitter account. More puzzling for me though is why the embedded XHTML metadata does not use Qualified Dublin Core! Let's crank up the interoperability, please!

Friday, 6 February 2009

Information seeking behaviour at Google: eye-tracking research

Anne Aula and Kerry Rodden have just published a posting on the Official Google Blog summarising some eye-tracking research they have been conducting on Google's 'Universal Search'. Both are active in information seeking behaviour and human-computer interaction research at Google and are well published within the related literature (e.g. JASIST, IPM, SIGIR, CHI, etc.).

The motivation behind their research was to evaluate the effect incorporation of thumbnail images and video within a research set has on user information seeking behaviour. Previous information retrieval eye-tracking research indicates that users scan results in order, scanning down their results until they reach a (potentially) relevant result, or until they decide to refine their search query or abandon the search. Aula and Rodden were concerned that the inclusion of thumbnail images might distract the "well-established order of result evaluation". Some comparative evaluation was therefore order of the day.
"We ran a series of eye-tracking studies where we compared how users scan the search results pages with and without thumbnail images. Our studies showed that the thumbnails did not strongly affect the order of scanning the results and seemed to make it easier for the participants to find the result they wanted."
A good finding for Google, of course; but most astonishing is the eye-tracking data. The speed with which users scanned result sets and the number of points on the interface they scanned was incredible. View the 'real time' clip below. A dot increasing in size denotes the length of time a user spent pausing at that specific point in the interface or result set. Some other interesting discoveries were made – the full posting is essential reading.