Wednesday, 25 March 2009

ASK conundrum revisited...again!

I posted a blog about Google's eye tracking research last month. I'm loathed to discuss Google again lest the ISG blog becomes known as the unofficial Google blog; however, the latest post on the Official Google Blog is worthy of some comment...

You might recall another post I made regarding search engine research and development, particularly in the area of information retrieval (IR) aids for users. In this posting I summarised Belkin's research and theories regarding the Anomalous State of Knowledge (ASK). Most of this and subsequent research has sought to introduce IR aids for the user so that they can better solve their ASK conundrum. This assistance varies but often takes the form of query expansion (in its various permutations), browsable subject trees to stimulate query formulation, relevance feedback, and so forth. Providing such tools in systems based on automatic indexing is difficult, but we noted that some search engines have introduced some effective retrieval aids, all designed to alleviate the ASK problem. For example, Yahoo! provides its search assist tool, Clusty provides related concept clusters, and Ask provides other similar tools. Their accuracy in IR varies widely, but overall they prove useful to user. Unfortunately, we also noted that Google provides few user aids comparable to those above, arguably relying more on its PageRank algorithm. Not any longer...

Today Google launched some interface functionality not dissimilar to Yahoo! search assist and Clusty. Their assistance provides some suggested related searches and some extra result summary text for particular results. Receiving this assistance depends on the nature of your query, so have a look at this canned search: 'communism in Russia'. This isn't bad and is better than nothing; but does it really measure up to the aids provided by competing search engines? Compare the results for these canned searches and the IR aids provided for the user by the systems we've discussed already:
Google's attempts appear quite pedestrian by comparison. Yahoo! and Clusty, for example, make their aids readily available so that the user can affect changes in their information seeking behaviour, but Google's tools are far less visible, less detailed, and offer far less functionality. Since a lot of research indicates that many users will not scroll below the 'golden triangle' (i.e. to the bottom of the first result set), it is entirely feasible to think that these 'related search' aids will go unnoticed by the disoriented information seeker.

It is good to see Google deploying user query aids and reacting to developments in other IR systems, but it appears that it will be some time before Google can be said to alleviate users' Anomalous State of Knowledge.

Tuesday, 17 March 2009

Who's going to teach our "stuff"? and who's going to learn it?

Is this the right forum for this? We all know of the imminent and unwelcome restructuring facing us. Where does the future lie for this discipline, or can we even define what our discipline is? We're constantly reminded now that our HE degrees are products, our students are customers, and so what are we, retailers? Compared to many other "products" in this HE marketplace our products are relatively unpopular despite the fact that there are fewer universities providing what we do compared to a decade ago.
I constantly struggle to explain what it is we do and can therefore understand why students have difficulty in placing it in context. Perhaps a business school isn't the right environment but then neither is a computing department, or it doesn't seem to be, and we don't fit into education or anywhere else.
What is the long term future for this discipline, whatever it is?

It's St. Patrick's day (not St. Paddy's, or St. Pat's or Paddy's day) so I'm off for Guinness in the local.

Thursday, 19 February 2009

Text is the new GUI?

We've got a software student ( from another University) working on his final year project at my day job. He is busy adding a 'speech' interface to our Laboratory tracking system. The idea is that while the scientist has their hands in the fume cupboard they don't want to be messing with a mouse or a keyboard why not engage with the computer by voice. This is after all how they did it in the old future in the movies.

Alas the student has got a bit sidetracked into the excitement of speech recognition and synthesis, in an attempt to get him to some sort of conclusion of his project I have suggested that most of the academic benefit of the project could be got by just having a text input and output (although useless in the fume cupboard). Once you go there it starts you thinking about how we interact with systems by voice. We are of course now used to listening to SatNav and some folk order their phones to phone the wife.

While pondering this and Georges earlier post of Twitter Library fees, I was thinking about an article about how fans have put together Twitter accounts of their favourite T.V. characters so that we can see when they are having a sandwich during the week. It made me wonder whether we might soon be engaging with various systems through the power of text rather than super 3D graphical interfaces.

This will be a shame as much of the design thought in web design and business application design has been about mice and windows more or less. In Liverpool this has led to some success for the hybrid 'programmer/graphic designer', perhaps if we are going to deal in a flow of text it will be the Hybrid 'programmer/DJ' or at least 'programmer/Linguist' who lead the way.

We surround ourselves in an increasing sense of a flow of consciousness through Twitter/Facebook etc. Surely this is going to include a Twitter from machines.

"Your fridge is enjoying a quiet day."
"Your car is worrying that it's service is due this week."
"Your door notes that fido is standing at it and wants to go out."

This will lead us to a wish to push application outputs into twitter like streams for our apps to respond to our own twitters. Text (or speech) may be the new GUI. Of course we will have to know a lot more about parsing and extracting meaning and identity from these streams of conciousness.

Thursday, 12 February 2009

FOAF and political social graphs

While catching up on some blogs I follow, I noticed that the Semantic Web-ite Ivan Herman posted comments regarding the US Congress SpaceBook – a US political answer to Facebook. He, in turn, was commenting on a blog made by the ProgrammableWeb – the website dedicated to keeping us informed of the latest web services, mashups, and Web 2.0 APIs.

From a mashup perspective, SpaceBook is pretty incredible, incorporating (so far) 11 different Web APIs. However, for me SpaceBook is interesting because it makes use of semantic data provided via FOAF and the microformat, XFN. To do this SpaceBook makes good use of the Google Social Graph API, which aims to harness such data to generate social graphs. The Social Graph API has been available for almost a year but has had quite a low profile until now. Says the API website:
"Google Search helps make this information more accessible and useful. If you take away the documents, you're left with the connections between people. Information about the public connections between people is really useful -- as a user, you might want to see who else you're connected to, and as a developer of social applications, you can provide better features for your users if you know who their public friends are. There hasn't been a good way to access this information. The Social Graph API now makes information about the public connections between people on the Web, expressed by XFN and FOAF markup and other publicly declared connections, easily available and useful for developers."
Bravo! This creates some neat connections. Unfortunately – and as Ivan Herman regrettably notes - the generated FOAF data is inserted into Hilary Clinton’s page as a page comment, rather than as a separate .rdf file or as RDFa. The FOAF file is also a little limited, but it does include links to her Twitter account. More puzzling for me though is why the embedded XHTML metadata does not use Qualified Dublin Core! Let's crank up the interoperability, please!

Friday, 6 February 2009

Information seeking behaviour at Google: eye-tracking research

Anne Aula and Kerry Rodden have just published a posting on the Official Google Blog summarising some eye-tracking research they have been conducting on Google's 'Universal Search'. Both are active in information seeking behaviour and human-computer interaction research at Google and are well published within the related literature (e.g. JASIST, IPM, SIGIR, CHI, etc.).

The motivation behind their research was to evaluate the effect incorporation of thumbnail images and video within a research set has on user information seeking behaviour. Previous information retrieval eye-tracking research indicates that users scan results in order, scanning down their results until they reach a (potentially) relevant result, or until they decide to refine their search query or abandon the search. Aula and Rodden were concerned that the inclusion of thumbnail images might distract the "well-established order of result evaluation". Some comparative evaluation was therefore order of the day.
"We ran a series of eye-tracking studies where we compared how users scan the search results pages with and without thumbnail images. Our studies showed that the thumbnails did not strongly affect the order of scanning the results and seemed to make it easier for the participants to find the result they wanted."
A good finding for Google, of course; but most astonishing is the eye-tracking data. The speed with which users scanned result sets and the number of points on the interface they scanned was incredible. View the 'real time' clip below. A dot increasing in size denotes the length of time a user spent pausing at that specific point in the interface or result set. Some other interesting discoveries were made – the full posting is essential reading.

Friday, 30 January 2009

Wikipedia: the new Knol?

Like many people I use Wikipedia quite regularly to check random facts. The strange aspect of this behaviour is that once I find the relevant fact, I have to immediately verify its provenance by conducting subsequent searches in order to find corroborative sources. It makes one wonder why one would use it in the first place.

Wikipedia continues to be plagued by a series of high profile malicious edits. Unfortunately, many of these edits aren't necessarily malicious. They are just wrong or inaccurate. There are probably hundreds of thousands of inaccurate Wikipedia articles, perhaps just as many hosting malicious edits; but it takes high profile gaffs to affect real change. On the day of Barack Obama's inauguration, Wikipedia reported the deaths of West Virginia's Robert Byrd and Edward Kennedy, who had collapsed during the inaugural lunch. Both reports were false.

This event appears to have compelled Jimmy Wales into being more proactive in improving the accuracy and reliability of Wikipedia. Under his proposals many future changes to articles would need to be approved by a group of vetted editors before being published. For me this news is interesting, particularly as it emerges barely two weeks after Google announced that the 100,000th Knol had been created on their 'authoritative and credible' answer to Wikipedia: Knol. Six or seven months ago the discussions focussed on how Knol was the new Wikipedia, now it appears as if Wikipedia might become the new Knol. How bizarre is that?!

Tightening the editing rules of Wikipedia has been on the agenda before and in 2007 this blog discussed how the German Wikipedia was conducting experiments which saw only trusted Wikipedians verifying changes to articles. So, will tightening the editing of Wikipedia make it the new Knol? The short answer is 'no'. Some existing Wikipedia editors can already exert authoritarian control over particular articles and can – in some cases – give the impression that they too have an axe to grind on particular topics. 'Wikipinochets' anyone? Moreover, Knol benefits from its "moderated collaboration" approach, with Knols being created by subject experts whose credentials have been verified. Wikipedia isn't going anywhere near this. Guardian columnist, Marcel Berlins, is probably right about Wikipedia when he states:
"I don't think there's a way of telling what proportion of Wikipedia entries are deficient, whether because of the writer's bias, mischief or lack of knowledge. It's clear that a significant number are questionable, sufficient to lead us to suspect all entries. But to do the right thing - vetting all contributors or contributions - would be impractical and hugely expensive. There is no easy solution. We many just have to accept that Wikipedia's undoubted usefulness comes at the price of occasional - perhaps frequent - inaccuracy. That is a sad conclusion to reach about an encyclopedia."
Oh well, back to verifying random facts found on Wikipedia.

Wednesday, 24 December 2008

SKOS-ifying Knowledge Organisation Systems: a continuing contradiction for the Semantic Web

A few days ago Ed Summers announced on his blog that he was shutting down lcsh.info. For those that don't know, lcsh.info was a Semantic Web demonstrator developed by Ed with the expressed purpose of illustrating how the Library of Congress Subject Headings (LCSH) could be represented and its structure harnessed using Simple Knowledge Organisation Systems (SKOS). In particular, Ed was keen to explore issues pertaining to Linked Data and the representation of concepts using URIs. He even hoped that the URIs used would be Cool URIs, linking eventually to a bona fide LCSH service were one ever to be released. Sadly, it was not to be... The reasons remain unclear but were presumably related to IPR. As the lcsh.info blog entry notes, Ed was compelled to remove it by the Library of Congress itself. The fact that he was the LC's resident Semantic Web buff probably didn't help matters, I'm sure.

SKOS falls within my area of interest and is an initiative of the Semantic Web Deployment Working Group. In brief, SKOS is an application of RDF and RDFS and is a series of evolving specifications and standards used to support the use of knowledge organisation systems (KOS) (e.g. information retrieval thesauri, classification schemes, subject heading systems, taxonomies or any other controlled vocabulary) within the framework of the Semantic Web. The Semantic Web is many things of course; but it is predicated upon the assumption that there exists communities of practices willing and able to create the necessary structured data (generally applications of RDF) to make it work. This might be metadata, or it might be an ontology, or it might be a KOS represented in SKOS. The resulting data can then be re-used, integrated, interconnected, queried and is open. When large communities of practice fail to contribute, the model breaks down.

There is a sense in which the Semantic Web has been designed to bring out the schizophrenic tendencies within some quarters of the LIS community. Whilst the majority of our community has embraced SKOS (and other related specifications), can appreciate the potential and actively contributes to the evolution of the standards, there is a small coterie that flirts with the technology whilst simultaneously shirking at the thought of exposing hitherto proprietary data. It's the 'lock down' versus 'openness' contradiction again.

In a previous research post I was involved with the High-Level Thesaurus (HILT) research project and continue my involvement in an consultative capacity. HILT continues to research and develop a terminology web service providing M2M access to a plethora of terminological data, including terminology mappings. Such terminological data can be incorporated into local systems to improve local searching functionality. Improvements might include, say, implementing a dynamic hierarchical subject browsing tree, or incorporating interactive query expansion techniques as part of the search interface, for example. An important - and the original motivation behind HILT - is to develop a 'terminology mapping server' capable of ameliorating the "limited terminological interoperability afforded between the federation of repositories, digital libraries and information services comprising the UK Joint Information Systems Committee (JISC) Information Environment" (Macgregor et al., 2007), thus enabling accurate federated subject-based information retrieval. This is a blog so detail will be avoided for now; but, in essence, HILT is an attempt to provide a terminology server in a mash-up context using open standards. To make the terminological data as usable as possible and to expose it to the Semantic Web, the data is modelled using SKOS.

But what happens to HILT when/if it becomes an operational service? Will its terminological innards be ripped out by the custodians of terminologies because they no longer want their data exposed, or will the ethos of the model be undermined as service administrators permit only HE institutions or charitable organisations from accessing the data? This isn't a concern for HILT yet; but it is one I anticipated several years ago. And the sad experience of lcsh.info illustrates that it's a very real concern.

Digital libraries, repositories and other information services have to decide where they want to be. This is a crossroads within a much bigger picture. Do they want their much needed data put to a good use on the Web, as some are doing (e.g. AGROVOC, GEMET, UKAT)? Or do they want alternative approaches to supplant them entirely (i.e. LCSH)? What's it gonna be, punks???

Monday, 8 December 2008

Wikipedia censorship: allusions to 'Smell the Glove'?

Another Wikipedia controversy rages, this time over censorship. Over the past two days, the Internet Watch Foundation informed some ISPs that an article pertaining to an album by the 'classic' German heavy metal band, Scorpions, may be illegal. Leaving aside the fact that Scorpions is one of many groups to have similar imagery on their record sleeves (the eponymous 1969 debut album by Blind Faith, Eric Clapton's supergroup, being another obvious example), am I the only person to notice the similarities with fictional rockumentary, This Is Spinal Tap?

Like metal, censorship is a heavy topic; but I thought this tenuous linkage with This Is Spinal Tap might be a welcome distraction from the usual blog postings, which are necessarily academic. Those of you familiar with said film might recall the controversy surrounding the proposed (tasteless) art work for Spinal Tap's new album (Smell the Glove), which in the end gets mothballed owing to its indecent nature. Getting into trouble over sleeve art is part and parcel of being in a heavy metal band it would seem! Enjoy the winter break, people!

Friday, 5 December 2008

Some general musings on tag clouds, resource discovery and pointless widgets...

The efficacy of collaborative tagging in information retrieval and resource discovery has undergone some discussion on this blog in the past. Despite emerging a good couple of years ago – and like many other Web 2.0 developments – collaborative tagging remains a topic of uncertainty; an area lacking sufficient evaluation and research. A creature of collaborative tagging which has similarly evaded adequate evaluation is the (seemingly ubiquitous!) 'tag cloud'. Invented by Flickr (Flickr tag cloud) and popularised by delicious (and aren't you glad they dropped the irritating full stops in their name and URL a few months ago?), tag clouds are everywhere; cluttering interfaces with their differently and irritatingly sized fonts.

Coincidentally, a series of tag cloud themed research papers were discussed at one of our recent ISG research group meetings. One of the papers under discussion (Sinclair & Cardew-Hall, 2008) conducted an experimental study comparing the usage and effectiveness of tag clouds with traditional search interface approaches to information retrieval. Their work is welcomed since it constitutes one of the few robust evaluations of tag clouds since they emerged several years ago.

One would hate to consider tag clouds as completely useless – and I have to admit to harbouring this thought. Fortunately, Sinclair and Cardew-Hall found tag clouds to be not entirely without merit. Whilst they are not conducive to precision retrieval and often conceal relevant resources, the authors found that users reported them useful for broad browsing and/or non-specific resource discovery. They were also found to be useful in characterising the subject nature of databases to be searched, thus aiding the information seeking process. The utility of tag clouds therefore remains confined to the search behaviours of inexperienced searchers and – as the authors conclude - cannot displace traditional search facilities or taxonomic browsing structures. As always, further research is required...

The only thing saving tag clouds from being completely useless is that they can occasionally assist you in finding something useful, perhaps serendipitously. What would be the point in having a tag cloud that didn't help you retrieve any information at all? Answer: There wouldn't be any point; but this doesn't stop some people. Recently we have witnessed the emergence of 'tag cloud generation' tools. Such tools generate tag clouds for Web pages, or text entered by the user. Wordle is one such example. They look nice and create interesting visualisations, but don't seem to do anything other than take a paragraph of text and increase the size of words based on frequency. (See the screen shot of a Wordle tag cloud for my home page research interests.)


OCLC have developed their very own tag cloud generator. Clearly, this widget has been created while developing their suite of nifty services, such as WorldCat, DeweyBrowser, FictionFinder, etc., so we must hold fire on the criticism. But unlike Wordle, this is something OCLC could make useful. For example, if I generate a tag cloud via this service, I expect to be able to click on a tag and immediately initiate a search on WorldCat, or a variety of OCLC services … or the Web generally! In line with good information retrieval practice, I also expect stopwords to be removed. In my example some of the largest tags are nonsense, such as "etc", "specifically", "use", etc. But I guess this is also a fundamental problem with tagging generally...

OCLC are also in a unique position in that they have access to numerous terminologies. This obviously cracks open the potential for cross-referencing tags with their terminological datasets so that only genuine controlled subject terms feature in the tag cloud, or productive linkages can be established between tags and controlled terms. This idea is almost as old as tagging itself but, again, has taken until recently to be investigated properly. Exploring the connections between tags and controlled vocabularies is something the EnTag project is exploring, a partner in which is OCLC. In particular, EnTag (Enhanced Tagging for Discovery) is exploring whether tag data, normally typified by its unstructured and uncontrolled nature, can be enhanced and rendered more useful by robust terminological data. The project finished a few months ago – and a final report is eagerly anticipated, particularly as my formative research group submitted a proposal to JISC but lost out to EnTag! C'est la vie!

Friday, 28 November 2008

Catching up with the old future of databases

We have been discussing in the group what we should be teaching on our Business Information Systems undergraduate course. Which is having a bit of a revamp. One area of discussion is about the areas of 'databases' which we teach mostly in the students second year and 'object oriented analysis and design' (but mostly analysis) which we teach in their final year.

When I used to teach database sections 10 years ago to business students we used to teach a history of:-
  1. file access
  2. hierarchical databases
  3. network databases
  4. relational databases
  5. object oriented databases.
Of course Object Oriented Databases hadn't happened in a big way back then. The surprising thing is that they haven't happened in a big way even now, they have spent 10 years being the next big thing. Meanwhile their close relatives UML based analysis and object oriented design and development have swept in from all directions. Strange that my notes from 10 years back now look like 'Space 1999', In their predictive powers. In my day job at Village we had a book on the subject 10 years back but a quick check of the book shelf shows we have long since recycled it.

Apparently believers and developers of Object Oriented Databases are just bemused by why everybody hasn't followed them into the promised land, particularly as much time is spent on 'Object Relational Mapping' technologies.

All round software development and architecture thinker and general purpose bearded Guru Martin Fowler, believes that the issue isn't to do with the general capabilities of the Object Oriented Databases but rather the fact that much integration in corporations occurs in the data layer not the business layer hence systems are dependent on standardised SQL approaches to integration. He suggests in his blog post on the subject that this shared database integration requirement has been holding back the march to the future of Object Oriented Databases. Creating extra inertia. Indeed, a confession, in my day job despite being Object Oriented N-Tier architecture developers by trade and conviction, when it came to tying our own timesheet system to our task management system we used database level triggers. It's a bit like the fact that there are better ways to do typing that using qwerty but we've all learnt to live with qwerty.

However with the movement towards using Web Services and SOA type architectures in effect making XML the linguq franca rather than SQL, Martin Fowler suggests that the field might start to loosen up. Although I wonder whether reporting is another issue. We produce some reports (in Crystal Reports and the equivalent) straight from our business objects, but other management reports really need to be produced straight off the (SQL) database. On occasions this is seperate from the main thrust of the application using a different technology stack.

Really the technology of the day in software development is Object Relational Mapping tools, ORMS. These try and hold the Object Oriented businesss layer to the Data Entity Oriented database layer. Such connections are relatively straightfoward in an unsophisticated design. My final year students currently angsting over a UML assignment will find that their Class Diagram is much the same as their Entity Relationship Diagram. But as you move deeper into doing things the Object way the two diverge. My Village colleague Ian Bufton and I have been discussing this in terms of lining up the two layers using either tools or code generation you can see some of his initial ponderings on his blog.

Luckily these types of contemplation are outside the scope of the things that our Information Systems students at the Business School, no doubt they have to worry about how to teach it at the Computer Science department.

Wednesday, 26 November 2008

'Gluing' searches with Yahoo!: part three in the search engine trilogy

There is plenty to comment on in the world of search engines at the moment. This post signifies the last in a series of discussions regarding search engines developments (or lack of!). (Part I; Part II)

As Google SearchWiki was unveiled, Yahoo! announced the wider release of Yahoo! Glue. Originally developed and tested at Yahoo! India, Yahoo! Glue now has a wider release – although it remains (perpetually?) in 'beta'. Glue is an attempt to aggregate disparate forms of information on a single results page in response to a single query. I suppose Glue is a functional demonstration of the ultimate mashup information retrieval tool. Glue assembles heterogeneous information from all over the Web, including text, news feeds, images, video, audio, etc. Search for the Beatles and you will get a results page listing Wikipedia definitions, LastFM tracks for your listening pleasure, news feeds, YouTube videos, etc.

In an ironic twist, Glue appears to be defying the dynamic ethos of Web 2.0. Glue searches are not created on the fly and only a limited number Glue searches are available at the moment (for example, no Liverpool!). Glue has the 'beta forever' mantra as the 'Get out of jail free card', of course. Still, Yahoo! informs us that:
"These pages are built using an algorithm that automatically places the most relevant modules on a page, giving you a visually rich, diverse page all about the topic in which you're interested."
Glue is also an example of Yahoo! exploring the social web in retrieval, harnessing as it does users' opinions on the accuracy of this algorithm (e.g. irrelevant or poorly ranked results can be 'flagged' as inappropriate or irrelevant).

Glue is - and will be - for the leisure user; the person falling into the 'popular search' category in search engines. These are the users submitting the simplest queries. The teenagers searching for 'Britney Spears', and the adults searching for 'Barack Obama' or 'Strictly Come Dancing'. The serious user (e.g. student, academic, knowledge worker, etc.) need not apply. I also have reservations over whether the summarisation of results is appropriate, and whether Glue can actually assemble disparate resources that are all relevant to a query. Check out these canned searches for Stephen Stills and Glasgow. In what way is Amsterdam a relation to Glasgow? And why the spurious news stories for Stephen Stills? Examining the text it is clear how it has been retrieved; but why when similar issues do not affect Yahoo! Search?

Like Google with SearchWiki, Yahoo! emphasise that Glue is not a replacements for Yahoo! Search; rather it's a "standalone experience":
"… Yahoo! Glue(TM) beta is not to replace the Yahoo! Search experience [...] We're always challenging ourselves to explore innovative new ways to deliver great experiences. Glue is one of those experiments, with a goal of giving users one more visual way to browse and discover new things from across the Web. We'll be working to expand the number of Glue pages, improve the experience and incorporate your feedback into future versions."
Very good. But making this dynamic and scalable should be atop the Glue 'to do' list. No Liverpool!

Monday, 24 November 2008

Wikifying search

This blog follows a series of other blogs pontificating about the efficacy of search engines in information retrieval. Over the weekend Google announced the release of Google SearchWiki. Google SearchWiki essentially allows users to customise searches by re-ranking, deleting, adding, and commenting on their results. This is personalised searching (see video below). As the Official Google Blog notes:
"With just a single click you can move the results you like to the top or add a new site. You can also write notes attached to a particular site and remove results that you don't feel belong."
The advantages of this are a little unclear at first; however, things become clearer when we learn that such changes can only be affected if you have an iGoogle account. Google have – quite understandably – been very specific about this aspect of SearchWiki. Search is their bread and butter; messing with the formula would be like dancing with the devil!

Google SearchWiki doesn't do anything further to address our Anomalous State of Knowledge (ASK), nor can I see myself using it, but it is an indication that Google is interested in better exploring the potential of social data to improve relevance feedback. Google will, of course, harvest vast amounts of data pertaining to users' information seeking behaviour which can then be channelled into improving their bread and butter. (And from my perspective, I would be interested to know how they analyse such data and affect changes in their PageRank algorithm). Their move also resonates with an increasing trend to support users in their Personal Information Management (PIM); to assist users in re-finding information they have previously located, or those frequently conducting the same searches over and over. It particularly reminds me of research undertaken by Bruce et al. (2004). For example, users increasingly chose not to bookmark a useful or interesting web page, but simply find it again – because they know they can. If you continually encounter information that is irrelevant to your area, re-rank it accordingly - so the SearchWiki ethos goes...

Perusing recent blogs it is clear that some consider this development to have business motivations. Technology guru and Wired magazine founder, John Battelle, thinks SearchWiki is an attempt to attract more users of iGoogle (which at the moment is small), whilst simultaneously rendering iGoogle the centre of users’ personal web universe. To my mind Google is always about business. PageRank is a great free-text searching tool, thus permitting huge market penetration. SearchWiki is simply another business tool which happens to offer some (vaguely?) useful functionality.

Tuesday, 7 October 2008

Search engines: solving the 'Anomalous State of Knowledge'

Information retrieval (IR) remains one of the most active areas of research within the information, computing and library science communities. It also remains one of the sexiest. The growth in information retrieval sex appeal has a clear correlation with the growth of the Web and the need for improvements in retrieval systems based on automatic indexing. No doubt the flurry of big name academics and Silicon Valley employees attending conferences such as SIGIR also adds glamour. Nevertheless, the allure of IR research has precipitated some of the best innovations in IR ever, as well as creating some of the most important search engines and business brands. Of course, asked to pick from a list their favourite search engine or brand, most would probably select Google.

The habitual use of Google by students (and by real people generally!) was discussed in a previous post and needn't be revisited here. Nevertheless, one of the most distressing aspects of Google (for me, at least!) is a recent malaise in its commitment to search. There have been some impressive innovations in a variety of search engines in a variety of areas. For example, Yahoo! is to better harness metadata and Semantic Web data on the Web. More interestingly though, some recent and impressive innovations in solving the 'ASK conundrum' is visible in a variety of search engines, but not in Google. Although Google always tell us that search is its bread and butter, is it spreading itself a little too thinly? Or - with a brand loyalty second to none and the robust PageRank algorithm deployed to good effect – is Google resting on its laurels?

In 1982 a young Nicholas J. Belkin spearheaded a series of seminal papers documenting various models of users' information needs in IR. These papers remain relevant today and are frequently cited. One of Belkin et al.'s central suppositions is that the user suffers from the so-called Anomalous State of Knowledge, which can be conveniently acronymized to 'ASK'. Their supposition can be summarised by the following quote from their JDoc paper:
"[P]eople who use IR systems do so because they have recognised an anomaly in their state of knowledge on some topic, but they are unable to specify precisely what is necessary to resolve that anomaly. ... Thus, we presume that it is unrealistic (in general) to ask the user of an IR system to say exactly what it is that she/he needs to know, since it is just the lack of that knowledge which has brought her/him to the system in the first place".
This astute deduction ushered in a branch of IR research that sought to improve retrieval by resolving the Anomalous State of Knowledge (e.g. providing the user with assistance in the query formulation process, helping users ‘fill in the blanks’ to improve recall (e.g. query expansion), etc.).

Last winter Yahoo! unveiled its 'Search Assist' facility (see screenshot above - search for 'united nations'), which provides a real time query formulation assistance to the user. Providing these facilities in systems based on metadata has always been possible owing the use of controlled vocabularies for indexing, the use of name authority files, and even content standards such as AACR2; but providing a similar level functionality with unstructured information is difficult – yet Yahoo! provide something ... and it can be useful and can actually help resolve the ASK conundrum!


Similarly, meta-search engine Clusty has provided its 'clustering' techniques for quite some time. These clusters group related concepts and are designed to aid in query formulation, but also to provide some level of relevance feedback to users (see screenshot above - search for 'George Macgregor'). Of course, these clusters can be a bit hit or miss but, again, they can improve retrieval and aid the user in query formulation. Similar developments can also be found in Ask. View this canned search, for example. What help does Google provide?

The bottom line is that some search engines are innovating endlessly and putting the fruits of a sexy research area to good use. These search engines are actually moving search forward. Can the same still be said of Google?

Friday, 29 August 2008

A conceptual model of e-learning: better studying effectiveness

My personal development has recently led me to explore and research the effectiveness of e-learning approaches to Higher Education (HE) teaching and learning. Since the late 1990s, e-learning has become a key focus of activity within pedagogical communities of practice (as well as those within information systems and LIS communities who often manage the necessary technology). HE is increasingly harnessing e-learning approaches to provide flexible course delivery models capable of meeting the needs of part-time study and lifelong learners. Of particular relevance, of course, is the Web, a mechanism highly conducive to disseminating knowledge and delivering a plethora of interactive learning activities (hence the role of informaticians).

The advantages of e-learning are frequently purported in the literature and are generally manifest in the Web itself. Such benefits include the ability to engage students in non-linear information access and synthesis; the availability of learning environments from any location and at any time; the ability for students to influence the level and pace of engagement with the learning process; and, increased opportunities for deploying disparate learning strategies, such as group discussion and problem-based or collaborative learning, as well as delivering interactive learning materials or learning objects. Various administrative and managerial benefits are also cited, such as cost savings over traditional methods and the relative ease with which teaching materials or courses can be revised.

Although flexible course delivery remains a principal motivating factor, the use of e-learning is largely predicated upon the assumption that it can facilitate improvements in student learning and can therefore be more effective than conventional techniques. This assumption is largely supported by theoretical arguments and underpins the large amounts of government and institutional investment in e-learning (e.g. JISC e-learning); yet, it is an assumption that is not entirely supported by the academic literature, containing as it does a growing body of indifferent evidence...

In 1983, Richard E. Clark from the University of Southern California conducted a series of meta-analyses investigating the influence of media on learning. His research found little evidence of any educational benefits and concluded that media were no more effective in teaching and learning than traditional teaching techniques. Said Clark:
"[E]lectronic media have revolutionised industry and we have understandable hopes that they would also benefit instruction".
Clark's paper was/is seminal and remains a common citation in those papers reporting indifferent e-learning effectiveness findings.

Is the same true of e-learning? Is there a similar assumption fuelling the gargantuan levels of e-learning investment? I feel safe in stating that such an assumption is endemic - and I can confirm this having worked briefly on a recent e-learning project. And I am in no way casting aspersions on my colleagues during this time, as I too held the very same assumption!!!

It is clear that evidence supporting the effectiveness of e-learning in HE teaching and learning remains unconvincing (e.g. Bernard et al.; Frederickson et al.). A number of comparative studies have arrived at indifferent conclusions and support the view that e-learning is at least as effective as traditional teaching methods, but not more effective (e.g. Abraham; Dutton et al.; Johnson et al.; Leung; Piccoli et al.). However, some of these studies exemplify a lack of methodological rigour (e.g. group self-selection) and many fail to control for some of the most basic variables hypothesised to influence effectiveness (e.g. social interaction, learner control, etc.). By contrast, those studies which have been more holistic in their methodological design have found e-learning to be more effective (e.g. Liu et al.; Hui et al.). These positive results could be attributed to the fact that e-learning, as an area of study, is maturing; bringing with it an improved understanding of the variables influencing e-learning effectiveness. Perhaps electronic media will "revolutionise" instruction after all?

Although such positive research tends to employ greater control over variables, such work fails to control for all the factors considered – both empirically and theoretically - to influence whether e-learning will be effective or not. Frederickson et al. have suggested that the theoretical understanding of e-learning has been exhausted and call for a greater emphasis on empirical research; yet it is precisely because a lack of theoretical understanding exists that invalid empirical studies have been designed. It is evident that the variables influencing e-learning effectiveness are multifarious and few researchers impose adequate controls or factor any of them into research designs. Such variables include: level of learner control; social interactivity; learning styles; e-learning system design; properties of learning objects used; system or interface usability; ICT and information literacy skills; and, the manner or degree to which information is managed within the e-learning environment itself (e.g. Information Architecture). From this perspective it can be concluded that no valid e-learning effectiveness research has ever been undertaken since no study has yet attempted to control for them all.

Motivated by this confusing scenario, and informed by the literature, it is possible to propose a rudimentary conceptual model of e-learning effectiveness (see diagram above) which I intend to develop and write up formally in the literature. The model attempts to improve our theoretical understanding of e-learning effectiveness and should aid researchers in comprehending the relevant variables and the manner in which they interact. It is anticipated that such a model will assist researchers in developing future evaluative studies which are both robust and holistic in design. It can therefore be hypothesised that using the model in evaluative studies will yield more positive e-learning effectiveness results.

Apologies this was such a lengthy posting, but does anyone have any thoughts on this or fancy working it up with me?

Friday, 8 August 2008

Where to next for social metadata? User Labor Markup Language?

The noughties will go down in history as a great decade for metadata, largely as a result of XML and RDF. Here are some highlights so far, off the top of my head: MARCXML, MODS, METS, MPEG-21 DIDL, IEEE LOM, FRBR, PREMIS. Even Dublin Core – an initiative born in the mid-1990s – has taken off in the noughties owing to its extensibility, variety of serialisations, growing number of application profiles and implementation contexts. Add to this other structured data, such as Semantic Web specifications (some of which are optimised for expressing indexing languages) like SKOS, OWL, FOAF, other RDF applications, and microformats. These are probably just a perplexing bunch of acronyms and jargon for most folk; but that's no reason to stop additions to the metadata acronym hall of fame quite yet...!

Spurred by social networking, so-called 'social metadata' has been emerging as key area of metadata development in recent years. For some, developments such as collaborative tagging are considered social metadata. To my mind – and those of others – social metadata is something altogether more structured, enabling interoperability, reuse and intelligence. Semantic Web specifications such as FOAF provide an excellent example of social metadata; a means of describing and graphing social networks, inferring and describing relationships between like-minded people, establishing trust networks, facilitating DataPortability, and so forth. However, social metadata is increasingly becoming concerned with modelling users' online social interactions in a number of ways (e.g. APML).

A recently launched specification which grabbed my attention is the User Labor Markup Language (ULML). ULML is described as an "open protocol for sharing the value of user's labor across the web" and embodies the notion that making such labour metric data more readily accessible and transparent is necessary to underpin the fragile business models of social networking services and applications. According to the ULML specification:
"User labor is the work that people put in to create, improve, and maintain their existence in social web. In more detail, user labor is the sum of all activities such as:
  • generating assets (e.g. user profiles, images, videos, blog posts),
  • creating metadata (e.g. tagging, voting, commenting etc.),
  • attracting traffic (e.g. incoming views, comments, favourites),
  • socializing with other people (e.g. number of friends, social influence)
in a social web service".
In essence then, ULML simply provides a means of modelling and sharing users' online social activities. ULML is structured much like RSS, with three major document elements (action, reaction and network). Check out the simple Flickr example below (referenced from the spec.). An XML editor screen dump is also included for good measure:

<?xml version="1.0" encoding="UTF-8"?>
<ulml version="0.1">
<channel>
<title>Flickr / arikan</title>
<link>http://fickr.com/photos/arikan</link>
<description>arikan's photos on Flickr.</description>
<pubDate>Thu, 06 Feb 2008 20:55:01 GMT</pubDate>
<user>arikan</user>
<memberSince>Thu, 01 Jun 2005 20:00:01 GMT</memberSince>
<record>
<actions>
<item name="photo" type="upload">852</item>
<item name="group" type="create">4</item>
<item name="photo" type="tag">1256</item>
<item name="photo" type="comment">200</item>
<item name="photo" type="favorite">32</item>
<item name="photo" type="flag">3</item>
<item name="group" type="join">12</item>
</actions>
<reactions>
<item name="photo" type="view">26984</item>
<item name="photo" type="comment">96</item>
<item name="photo" type="favorite">25</item>
</reactions>
<network>
<item name="connection">125</item>
<item name="density">0.167</item>
<item name="betweenness">0.102</item>
<item name="closeness">0.600</item>
</network>
<pubDate>Thu, 06 Feb 2008 20:55:01 GMT</pubDate>
</record>
</channel>
</ulml>


The tag properties within the action and reaction elements are all pretty self-explanatory. Within the network element "connection" denotes the number of friend connections, "density" denotes the number of connections divided by the total number of all possible connections, "closeness" denotes the average distance of that user to all their friends, and "betweenness" the "probability that the [user] lies on the shortest path between any two other persons".

Although the specification is couched in a lot of labour theory jargon, ULML is quite a funky idea and is a relatively simple thing to implement at an application level. With the relevant privacy safeguards in place, service providers could make ULML files publicly available, thus better enabling them and other providers to understand users' behaviour via a series of common metrics. This, in turn, could facilitate improved systems design and personalisation since previous user expectations could be interpreted through ULML analysis. Authors of the specification also suggest that a ULML document constitutes an online curriculum vitae of users' social web experience. It provides a synopsis of user activity and work experience. It is, in essence, evidence of how they perform within social web contexts. Say the authors:
"...a ULML document is a tool for users to communicate with the web services upfront and to negotiate on how they will be rewarded in return for their labour within the service".
This latter concept is significant since – as we have discussed before – such Web 2.0 services rely on social activity (i.e. labour) to make their services useful in the first place; but such activity is ultimately necessary to make them economically viable.

Clearly, if implemented, ULML would be automatically generated metadata. It therefore doesn't really relate to the positive metadata developments documented here before, or the dark art itself; however, it is a further recognition that with structured data there lies deductive and inferential power.