Tuesday, 22 June 2010

Goulash all round: Linked Data at NSZL

I meant to blog about this as soon as the news emerged in mid-April but University bureaucracy and research project demands prevented it: Adam Horvath (Director of Informatics) at the The National Széchényi Library (NSZL) (or National Library of Hungary, if you prefer) announced on the Semantic Web Linking Open Data Project email list that the NSZL have exposed their entire OPAC and digital library as Linked Data - that's correct, their entire OPAC and digital library has been published as Linked Data. This includes corresponding authority data, with all nodes represented using Cool URIs.

The RDF vocabularies used include Dublin Core RDF for bibliographic metadata, SKOS for subject indexing (in a variety of terminologies) and FOAF for name authority data. Incredible! Not only that, the FOAF descriptions include mapped owl:sameAs statements to corresponding dbpedia URIs. For example, here is FOAF data pertaining to Hungarian novelist, Jókai Mór:

<?xml version="1.0"?>
<rdf:RDF
xmlns:dbpedia="http://dbpedia.org/property/"
xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
xmlns:skos="http://www.w3.org/2004/02/skos/core#"
xmlns:foaf="http://xmlns.com/foaf/0.1/"
xmlns="http://web.resource.org/cc/"
xmlns:owl="http://www.w3.org/2002/07/owl#"
xmlns:dc="http://purl.org/dc/elements/1.1/"
xmlns:zs="http://www.loc.gov/zing/srw/">
<foaf:Person rdf:about="http://nektar.oszk.hu/resource/auth/33589">
<dbpedia:deathYear>1904</dbpedia:deathYear>
<dbpedia:birthYear>1825</dbpedia:birthYear>
<foaf:familyName>Jókai</foaf:familyName>
<foaf:givenName>Mór</foaf:givenName>
<foaf:name>Jókai Mór (1825-1904)</foaf:name>
<foaf:name>Mór Jókai</foaf:name>
<foaf:name>Jókai Mór</foaf:name>
<owl:sameAs rdf:resource="http://dbpedia.org/resource/M%C3%B3r_J%C3%B3kai"/>
</foaf:Person>
</rdf:RDF>


Visit the above noted dbpedia data for fun.

Rich SKOS data is also available for a local information retrieval thesaurus. Follow this link for an example of the skos:prefLabel ' magyar irodalom'.

It's a herculean effort from the NSZL which must be commended. And before the Germans did it too! Goulash all round to celebrate - and a photograph of the Hungarian Parliament, methinks.

Tuesday, 30 March 2010

Social media and the organic farmer

The latest Food Programme was broadcast yesterday by BBC Radio 4 and made for some interesting listening. (Listen again at iPlayer.) In it Sheila Dillon visited the Food and Drink Expo 2010 at the Birmingham NEC and, rather than discussing the food, her focus was encapsulated in the programme slogan, 'To Tweet or not to Tweet', which was also the name of a panel debate at the Expo. Twitter was not the principal programme focus though. The programme explored social media generally and its use by small farmers and food producers to communicate with customers.

There were some interesting discussions about the fact that online grocery sales grew by 15% last year (three times more than 'traditional' grocery sales) and the role of the web and social media in the disintermediation of supermarkets as the principal means of getting 'artisan foods' to market. Some great success stories were discussed, such as Rude Health and the recently launched Virtual Farmers' Market. However, the familiar problem which the programme highlighted – and the problem which has motivated this blog posting – is the issue of measuring the impact and effectiveness of social media as a marketing tool. Most small businesses had little idea how effective their social media 'strategies' had been and, by the sounds of it, many are randomly tweeting, blogging and setting up Facebook groups in a vein attempt to gain market traction. One commentator (Philip Lynch, Director of Media Evaluation) from Kantar Media Intelligence spoke about tracking "text footprints" left by users on the social web which can then be quantified to determine the level support for a particular product or supplier. He didn't say much more than that, probably because Kantar's own techniques for measuring impact are a form of intellectual property. It sounds interesting though and I would be keen to see it action.

But the whole reason for the 'To Tweet or not to Tweet' discussion in the first place was to explore the opportunities to be gleaned by 'artisan food' producers using social media. These are traditionally small businesses with few capital resources and for whom social media presents a free opportunity to reach potential customers. Yet, the underlying (but barely articulated) theme of many discussions on the Food Programme was that serious investment is required for a social media strategy to be effective. The technology is free to use but it involves staff resource to develop a suitable strategy, and a staff resource with the communications and technical knowledge. On top of all this, small businesses want to be able to observe the impact of their investment on sales and market penetration. Thus, in the end, it requires outfits like Kantor to orchestrate a halfway effective social media strategy, maintain it, and to measure it. Anything short of this will not necessarily help drive sales and may be wholly ineffective. (I can see how social media aficionado Keith Thompson arrived at a name for his blog – any thoughts on this stuff, Keith?) The question therefore presents itself: Are food artisans, or any small business for that matter, being suckered by the false promise of free social media?

Of course, most of the above is predicated upon the assumption that people like me will continue to use social media such as Facebook; but while it continues to update its privacy policy, as it suggested this week on its blog that it will, I will be leaving social media altogether. 'Opting in' for a basic level of privacy should not be necessary.

Monday, 29 March 2010

Students' information literacy: three events collide with cosmic significance...

Three random – but related – events collided last week, as if part of some cosmic information literacy solar system...

Firstly, I completed marking student submissions for Business Information Management (LBSIS1036). This is a level one module which introduces web technologies to students; but it is also a module which introduces information literacy skills. These skills are tested in an in-lab assessment in which students can demonstrate their ability to critically evaluate information, ascertain provenance, IPR, etc. To assist them the students are introduced to evaluation methodologies in the sessions preceding the assessment which they can use to test the provenance of information sources found on the 'surface web'.

Students' performance in the assessment was patchy. Those students that invested a small amount of time studying found themselves with marks within the 2:1 to First range; but sadly most didn't invest the preparation time and found themselves in the doldrums, or failing altogether. What was most revealing about their performance was the fact that – despite several taught sessions outlining appropriate information evaluation methodologies – a large proportion of students informed me in their manuscripts that their decision to select a resource was not because i t fulfilled particular aspects of their evaluation criteria, but because the resource featured in the top five results within Google and therefore must be reliable. Indeed, the evaluation criteria were dismissed by many students in favour of the perceived reliability of Google's PageRank to provide a resource which is accurate, authoritative, objective, current, and with appropriate coverage. Said one student in response to 'Please describe the evaluation criteria used to assess the provenance of the resource selected': "The reason I selected this resource is that it features within the top five results on Google and therefore is a trustworthy source of information".

Aside from the fact these students completely missed the point of the assessment and clearly didn't learn anything from Chris Taylor or me, it strikes fear in the heart of a man that these students will continue their academic studies (and perhaps their post-university life) without the most basic information literacy skills. It's a depressing thought if one dwells on it for long enough. On the positive side, only one student used Wikipedia...which leads me to the next cosmic event...

Last week I was doing my periodic 'catch up' on some recent research literature. This normally entails scanning my RSS feeds for recently published papers in the journals and flicking through the pages of the recent issues of the Journal of the American Society for Information Science and Technology (JASIST). A paper published in JASIST at the tail end of 2009 caught my eye: 'How and Why Do College Students Use Wikipedia?' by Sook Lim which, compared to the hyper scientific paper titles such as 'A relation between h-index and impact factor in the power-law model' or 'Exploiting corpus-related ontologies for conceptualizing document corpora' (another interesting paper), sounds quite magazine-like. Lim investigated and analysed data on students' perceptions, uses of, and motivations for using Wikipedia in order to better understand student information seeking behaviour. She employed frameworks from social cognitive theory and 'uses and gratification' literature. Her findings are too detailed to summarise here. Suffice to say, Lim found many students to use Wikipedia for academic purposes, but not in their academic work; rather, students used Wikipedia to check facts and figures quickly, or to glean quick background information so that they could better direct their studying. In fact, although students found Wikipedia to be useful for fact checking, etc., their perceptions of its information quality were not high at all. Students knew it to be a suspect source and were sceptical when using it.

After the A&E experience of marking the LBSIS1036 submissions, Lim's results were fantastic news and my spirits were lifted immediately. Students are more discerning than we give them credit for, I thought to myself. Fantastic! 'Information Armageddon' doesn't await Generation Y after all. Imagine my disappointment the following morning when I boarded a train to Liverpool Central to find myself seated next to four students. It was here that I would experience my third cosmic event. Gazing out the train window as the sun was rising over Bootle docks and the majesty of its containerisation, I couldn't help but listen to the students as they were discussing an assignment which they had all completed and were on their journey to submit. The discussion followed the usual format, e.g. "What did you write in your essay?" "How did you structure yours?", etc. It then emerged that all four of them had used Wikipedia as the principal source for their essay and that they simply copied and pasted passages verbatim. In fact, one student remarked, "The lecturer might get suspicious if you copy it directly, so all I do is change the order of any bullet points and paragraphs. I change some of the words used too". (!!!!!!!!!!)

My hope would be that these students get caught cheating because, even without using Turnitin, catching students cheating with sources such as Wikipedia is easy peasy. But a bigger question is whether information literacy instruction is a futile pursuit? Will instant gratification always prevail?

Image: Polaroidmemories (Flickr), CreativeCommons Attribution-Non-Commercial-Share Alike 2.0 Generic

Friday, 26 March 2010

Here cometh the pay wall: the countdown begins...

So the Times Online will be charging for news content from June 2010... News International is just one of several news organisations (some big, e.g. NY Times, ABC, etc.) which have plans to erect pay walls and develop subscription based models imminently. Love him or loathe him (well, most people loathe him), you have to respect Rupert Murdoch's pit-bull instincts in leading a growing number of content providers to erect – or at least consider the erection – of pay walls. Murdoch knows that most of the content industry wants to see the proliferation of pay walls. In fact, pay walls are their only saviour. Certain death awaits them otherwise. James Harding of The Times said as much in an interview today: "It [charging for Times Online access] is less of a risk than continuing to do what we are currently doing".

The trouble is that few yet have the gumption to do it. Murdoch, I suspect, is one of several who thinks that once there is a critical mass of high profile content providers implementing pay walls then there will be deluge of others. And I think he is probably correct in this assumption. After all, subscription can actually work. The FT and Wall Street Journal have successfully had subscription models for years (although they admittedly provide an indispensible service to readers in the financial and business sectors). An additional benefit of pay wall proliferation will be the simultaneous decline of news aggregators (which Murdoch has been particularly vexed about recently) and 'citizen journalists', both of which have contributed to the ineffectiveness of advertising as a business model for online newspapers. The truth is that the future of good journalism depends on the success of these subscription-based business models; but the success of this also has implications for other content providers or Internet services experiencing similar problems, social networking services being a prime example.

If you take the time to peruse the page created by BBC News to collect user comments on this story, a depressingly long slew of comments can be found in which it becomes clear that most users (not all, it should be noted) have a huge difficulty with subscription models or simply do not understand what the business problem is. And largely this is down to the fact that most ordinary people think:
  1. That content providers of all types, not just newspapers, are a bunch of rip-off merchants who are dissatisfied with their lot in the digital sphere;
  2. That content providers generate abnormal profits from advertising revenue and that their businesses are based on robust business models, and
  3. That free content, aggregation and 'citizen journalists' fulfil their news or content needs admirably and that high quality journalism is therefore superfluous.
My opinion, for what it is worth, is that most users simply don't recognise that businesses require business models, nor do they realise that many of the services they use on the web day in and day out are either unprofitable or are losing large amounts of money. It's always good to receive something for free; however, someone somewhere has to pay. This reality is inescapable. Traditionally this has been advertisers, partly because Google have been good at it; but what do you do when advertising doesn't bring home the bacon??? Many users dismiss the implementation of pay walls by telling us that they will get their news from a free sources or blogs. The reliance on free or citizen journalism is disappointing, but more than that it is dangerous for democracy. Such sources simply do not have the resources (financial or intellectual) to deliver high quality, reliable news. They don't have the training, or the new connections, or the international correspondents, or the access to the information required, nor do they operate within recognised ethical boundaries or present facts and stories objectively and with appropriate sources or evidence. Often they are motivated more by communicating with like minded readers, perpetuating gossip or untruths.

The real truth of course is that newspapers are losing tremendous amounts of money. It seems to be unfashionable to say it – even the great but struggling Guardian chokes on these words – but free can no longer continue. Newspapers across the world have restructured and reinvented themselves to cope with the digital world. But there is only so much rearranging of the Titantic's deck chairs that can occur. The bottom line is that advertising as a business model isn't a business model. (See this, this and this for previous musings on this blog). Facebook is set to be 'cash flow positive' for the first time this financial year. No-one knows how much profit it will generate, although economic analysts suspect it will be small. What does this say about the viability of advertising as a revenue stream when a service with over 500 million users can barely cover costs? But what are Facebook to do? Chris Taylor conducted an unscientific survey with some International MBA students last week, all of whom reported positively on their continued use of Facebook to connect with family and friends at home and within Liverpool Business School. The question was: Would you be willing to pay £3 per annum to access Facebook? The response was unanimous: 'No'.

Sigh.

Thursday, 25 March 2010

How much software is there in Liverpool and is it enough to keep me interested

I am worrying about 'S' which I'll define here as the quantity of commercial software source code under maintenance and 'delta S' how this figure is changing. And by implication what we should be doing to make 'S' and 'delta-S' bigger. I'd rather the government worried about this rather than subsidising super fast YouTube to cottage dwellers.

In particular I am thinking about my home market in Liverpool where I am hoping to continue to carve some kind of career in my day job. If there is not enough 'S' to keep me going for another 25 years I'm going to get bored, poor or a job at McDonalds.

'S' maintenance is only a relatively minor destination for our business information systems students (still the students with the highest exit salary in the business school I am told please send yourself and your children BIS Course), but it is important to me.

Back in the day job we are working on developing LabCom a business to business tracking system for chemical samples and their results. One of the things that appeals to my mind is the fact that it is building a machine that is making things happen. I like to see how many samples are processed on it a year. Sad I know. We are delivering various new modules that will hopefully allow it all to grow. However thus far the project is not really big enough to fund the level of technical development and architecture expertise we would like to deploy, not enough 'S' to on its own maintain and fund high level development capabilities. The alternative for software developers such as my team is to engage in shorter term development consultancy forays, but for these to be of sustained technical interest they have probably got to add up to 100K or more and alas we have not worked out how to regularly corner such jobs.

With few software companies in Liverpool I wonder how this translates into the bigger picture and whether we can measure it.

There is definitely some 'S' in Liverpool, I did some work a couple of years back looking at software architecture with my friends at New Mind who are a national leader in destination management services and have a big crunching bit of software behind it. Angel Solutions are another company with a national footprint, this time in the education sector, backed by source code controlled in Liverpool. I have come across only two or three others over the years although no doubt there are some hiding. For someone trying to make a career out of having the skills to understand and develop big software this lack of available candidates in Liverpool might be a bit of a problem. A bit like getting an advanced mountaineering certificate in Norfolk ( see news Buscuit).

With this in mind I wondered is there more or less 'S' under management in Liverpool than elsewhere. Is this really the Norfolk of mountaineering.

How can we know. There are some publicly available records, we could dig up finances of software companies and similar, although most of these companies (including my own) are principally guns for hire engaging in consultancy and development services.

Liverpool's Software/New Media industry has a number of great companies such as Liverpool Business School Graduate led Mando Group and Trinity Mirror owned Ripple Effect and the only big player Strategic System Solutions. However as I understand it these are service delivery and consultancy companies not software product development companies they make their money through expertese, they contain a relatively low proportion of 'S'. Probably the largest block of software under management will be in the IT departments no doubt they have a bit of 'S'.

So how could we weigh this here in Liverpool or elsewhere. How much software is there. Lots of small companies such as my own have part of their income from owned source code IP 'S'. There are the few larger ones. So we could try to get a list and determine what proportion of income is generated from these outfits based on public records and a little inside knowledge. We could perhaps measure the number of software developers deployed or the traditional measure of SLOC (Source Lines Of Code). I've in the past looked at variants on Mark II Function Point Analysis you can find out a little about this on the United Kingdom Software Metrics Association website . In my commercial world I’m interested in estimating cost hence toying with these methods while we ponder what we can get away with charging and is it more than the our estimated (guessed) cost. In this regional context I’m interested in whether we can measure how much value there is lurking to give a figure for 'S'.

Imagine we did measure that the quantity of commercial code under management in Liverpool was I’ll call it 'Liver-S', I then want to know how this compares to Manchester's 'Manc-S' (I suspect unfavourably) and perhaps more importantly from a career and commercial point of view how it compares to last year. Is 'Liver-S' getting bigger (what is 'delta Liver-S').
The importance of 'S' and 'delta S' is about whether we are maintaining enough work to maintain or indeed develop a capacity to ‘do big software’ in the local economy. Without which to be honest I’m going to get bored.

Any masters/mba students stuck for a bit of an assignment or even better funding bodies wanting to help me answer this question please drop me a line.

If there is not enough 'Liver-S' in the future at least I'll be able to sit at home with my super fast broadband.

Tuesday, 2 February 2010

The 'real' iPad

I am now utterly exhausted with all the discussion and comment surrounding the new Apple iPad. Phew! The amusing video below therefore offers a welcome break from the hyperbole.

The video, publicised by BoingBoing, features one of Liverpool's funniest chaps: Peter Serafinowicz. Serafinowicz has managed to release an iPad parody (days after its launch by Apple) in order to plug the release of his latest DVD. He's quick off the mark. Serafinowicz has often included Apple parodies in his shows, such as the iToilet and the MacTini. This follows a similar format and is similarly amusing. Enjoy!

The iPad - watch more funny videos

Tuesday, 26 January 2010

Renaissance of thesaurus-enhanced information retrieval?

As my students in BSNIM3034 Content Management have been learning recently, semantics play such a huge role in the level of recall and precision capable of being achieved in an information retrieval system. Put simply, computers are great at interpreting syntax but are far too dumb to understand semantics or the intricacies of human language. This has historically been – and currently remains – the trump card of metadata proponents, and it is something the Semantic Web is attempting to resolve with its use of structured data too. The creation of metadata involves a human; a cataloguer who performs a 'conceptual analysis' of the information object in question to determine its 'aboutness'. They then translate this into the concepts prescribed in a controlled vocabulary or encoding scheme (e.g. taxonomy, thesaurus, etc.) and create other forms of descriptive and administrative metadata. All this improves recall and precision (e.g. conceptually similar items are retrieved and conceptually dissimilar items are suppressed).

As good as they are these days, retrieval systems based on automatic indexing (i.e. most web search engines, including Google, Bing, Yahoo!, etc.) suffer from the 'syntax problem'. They provide what appears to be high recall coupled with poor precision. This is the nature of search the engine beast. Conceptually similar items are ignored because such systems are unable to tell that 'Haematobia irritans' is a synonym of 'horn flies' or that 'java' is a term fraught with homonymy (e.g. 'java' the programming language, 'java' the island within the Indonesian archipelago and 'java' the coffee, and so forth). All of the aforementioned contributes to arguably the biggest problem for user searching: query formulation. Search engines suffer from the added lack of any structured browsing of, say, resource subjects, titles, etc. to assist users in query formulation.

This blog has discussed query formulation in the search process at various times (see this for example). The selection of search terms for query formulation remains one of the most difficult stages in users' information retrieval process. The huge body of research relating to information seeking behaviour, information retrieval, relevance feedback, and human-computer interaction attests to this. One of the techniques used to assist users in query formulation is thesaurus assisted searching and/or query expansion. Such techniques are not particularly new and are often used in search systems successfully (see Ali Shiri's JASIST paper from 2006).

Last week, however, Google announced adjustments to their search service. This adjustment is particularly significant because it is an attempt to control for synonyms. Their approach is based on 'contextual language analysis' rather than the use of information retrieval thesauri. The blog reads:
"Our systems analyze petabytes of web documents and historical search data to build an intricate understanding of what words can mean in different contexts [...] Our synonyms system is the result of more than five years of research within our web search ranking team. We constantly monitor the quality of the system, but recently we made a special effort to analyze synonyms impact and quality."
Firstly, this is certainly positive news. Synonyms – as noted above – are a well known phenomenon which has blighted the effectiveness of automatic indexing in retrieval. But on the negative side – and not to belittle Google's efforts as they are dealing with unstructured data - Google are only dealing with single words. 'Song lyrics' and 'song words', or 'homocide' and 'murder' (examples provided from Google on their blog posting) They are dealing with words in a Roget's Thesaurus sense, rather than compound terms in an information retrieval thesaurus sense – and it is the latter which will ultimately be more useful in improving recall and precision. This is, after all, why information retrieval thesauri have historically been used in searching.

More interesting will be Google's exploration of homonymous terms. Homonyms are more complex that synonyms and are, perhaps for the foreseeable future, an intractable problem?

Friday, 15 January 2010

Death of the book salesman...

Shopping in Glasgow prior to Christmas was a sad time. Borders, which occupied what is reputed to be the most expensive retail space in Glasgow (the old Royal Bank of Scotland building), announced that it was in administration and was flogging all stock in a gargantuan clearance sale. Borders had become an institution since it opened on Buchannan Street in 1997 (I think) and I'm sure branches in other cities were similarly iconic and located at city centre hot-spots, the London Oxford Street branch being another prime example. It was a great place to meet friends before heading out for dinner or drinks; perusing the amazing magazine or newspaper selection, or browsing the books or music. Of course, I stopped buying books there years ago because the genre classification they used made it impossible to find anything; but it nevertheless occupied a special place in my heart...

The official line is that Borders fell victim to the current economic climate, although it was a complicated concatenation of economic circumstances, including aggressive competition from online retailers and particularly supermarkets (those supermarkets again – a £3 copy of the latest Jordan autobiography anyone?), a sales downturn and, finally, a lack of credit from suppliers. Waterstone's remains the only national bookseller but today was responsible for a decline in the share price of HMV as their Chief Executive tries to administer 'bookshop CPR' (i.e. let's make our stores more cosy). Can we expect the closure of it too in the foreseeable future? That would be extremely depressing...

Of course it's all depressing news; but one can't help thinking that the demise of super-selling bookshops was a quagmire of their own making. The Net Book Agreement (NBA) – the 100 year long (almost) price fixing of books which collapsed in 1995 – was precipitated by Waterstone's in the first place. And it was precisely the collapse of the NBA which enticed Borders to the UK and enabled Amazon to establish UK operations. Both retailers would not have been able to operate with the NBA still in operation (remember the big Amazon book discounts in the late 1990s?). I suppose none of the book retailers anticipated the level of competitive aggression they had unleashed, particularly from supermarkets. Although I think some naivety played a part...

Many years ago I recall enjoying a talk delivered by the Deputy Chairman of John Smith & Son, Willie Anderson. Despite being the oldest bookseller in the English speaking world, John Smith moved off the high street many years ago. You are now most likely to encounter them as the university campus bookseller. During his talk Willie made an interesting point about the lack of business sense in the book selling industry; that the NBA had made all book sellers blind to conventional business practice or simple economic principles such as the laws of supply and demand. Said Willie (as best as I can remember! It was well over 10 years ago!):
"Harry Potter and the Chamber of Secrets was anticipated to be a best seller and an extremely popular title. We [John Smith & Son] had large pre-orders from customers. Yet, virtually all other booksellers were slashing prices and offering ridiculous pre-order discounts on an item which commanded a high price. At John Smith we didn't offer any discounts and we sold every copy at full price, precisely because demand was high. This is normal business practice, but most book retailers appear to be oblivious to this. Book retailers have a lot to learn about competition because they have been protected from it for so long. The industry needs to learn quickly otherwise it will suffer economic difficulties in the future".
What happens now then? There is certainly money to be made in book selling, particularly with the decline of 'good' stockists. There are more books bought now than at any time in history. Perhaps the time is ripe for a renaissance in the classic independent bookshop, of which Reid of Liverpool is archetypal? Supermarkets do not - I think - occupy the same business space as such book sellers and thus allowing the independent retailer to thrive. There wouldn't be any Costa or Starbucks, nor would it occupy a prime retail site, but I think we'd be all the better for it.

(Photo: Laura-Elizabeth, Flickr, Creative Commons)

Tuesday, 17 November 2009

Getting into technical debt in a recession?

I like this succinct quote from TalkTalk CIO David Cooper about the IT systems in newly acquired Tiscali.

"In addition, Tiscali faced some IT issues in the past and worked pragmatically to fix them, resulting in some discontinuities between systems – we are repairing this now," www.computing.co.uk/computing/analysis/2252848/ringing-changes-talktalk-4893004

There are, no doubt, times to go for the cheapest, quickest most 'pragmatic' solution to getting systems to mesh together. But in getting to a solution for least cost there can be an accumulation of 'Technical Debt'. I know that one of my clients in my day job, keeps a measure of technical debt being accumulated in development projects. But it is difficult to persuade an organisation to contemplate such debt, let alone put it on the balance sheet.

It's never easy either to argue for spending money on technical debt when you can have new shiny functionality baubles. What we find we do at Village Software when working on clients projects is try to improve things as we go along. Most clients would not be happy if we suggested they spent 10’s of thousands refactoring a system with little or no functional gain. Hence I suspect we often do this on the cheap out of a possibly misplaced or at least poorly negotiated sense of professionalism. We have recently spent a five figure sum refactoring our Lab Solution to improve it under the hood, I've got to say this hurts and we are now on a functionality campaign.

I wonder what the effect of the current recession is on technical debt. In principle resources to deal with it are cheaper than in a boom time. However the need to gain the cost reducing, innovation gaining benefits of Business Information Systems at lower investments will surely lead to an increase in technical debt across the economy. Perhaps the UK is accumulating billions of technical debt in the public and private sector to match the vast national debt in the public sector and the balance sheet retrenchments in the private sector.

The term 'Technical Debt' by the way was coined by Ward Cunningham as a useful allegory. One of the great thinkers in current software development Martin Fowler describes it on his Wiki martinfowler.com/bliki/TechnicalDebt.html. He explains that getting things done, in a way TalkTalk's David Cooper describes generously as pragmatically, is like borrowing money, you have to start paying interest, eventually you have to pay it back along with the principal. Ward Cunningham has a neat little 4 point plan referring to technical debt in software development. (For those not familiar with the term, refactoring is the practice of improving software code quality without adding functionality). Cunningham describes technical debt on his Wiki (all these guys have wikis) www.c2.com/cgi/wiki?ComplexityAsDebt :-
  • Skipping design is like borrowing money.
  • Refactoring is like repaying principal.
  • Slower development due to complexity is like paying interest.
  • When the whole project caves in under the mess, is that like when the big guys come round and slam your hands in the car door for not paying up?

Others describe it by the allegory of lactic acid building up on your muscles during a run. In the wider world of Commercial and Government ICT we can expect a build up of such debt. For businesses without a strategic plan for ICT, the resources to deliver a plan or whose plan is to accumulate technical debt there is going to be a backlog. Alas 'Technical Debt' will not be appearing on balance sheets soon if indeed there were a way to measure it. I would be baffled if a student asked how to go about measuring technical debt in his company.

I have a feeling that the Information Systems people, like me, will somehow get the blame for letting such debt accumulate.

Life can be unfair, but it is indoor work no heavy lifting.

Friday, 13 November 2009

An interesting article about search engines...again...

Victor Keegan (Guardian Technology journalist) published an interesting column yesterday on the current state of search engines. The column entitled, "Why I am searching beyond Google", is an interesting discussion which picks up on something that has been discussed a lot on this blog: the fact that Google really isn't that good any more. There are dozens of search engines out there which offer the user greater functionality and/or search data which Google ignores or can't be bothered indexing. Yahoo! and Bing are mentioned by Keegan, but leapfish, monitter and duckduckgo are also discussed.

Keegan also comments on the destructive monopoly that Google has within search:
"If you were to do a blind tasting of Google with Yahoo, Bing or others, you would be pushed to tell them apart. Google's power is no longer as a good search engine but as a brand and an increasingly pervasive one. Google hasn't been my default search for ages but I am irresistibly drawn to it because it is embedded on virtually every page I go to and, as a big user of other Google services (documents, videos, Reader, maps), I don't navigate to Google search, it navigates to me."
This is where Google's dominance is starting to become a problem. Competition is no longer fair. There are now several major search engines which are, in many ways, better than Google; yet, this is not reflected in their market share, partly because the search market is now so skewed in Google's favour. As Keegan notes, Google comes to him, not the other way round.

In a concurrent development, WolframAlpha is to be incorporated into Bing to augment Bing's results in areas such as nutrition, health and mathematics. Will we see Google incorporate structured data from Google Squared into their universal search soon?

I realise that this is yet another blog posting about either, a) Google, or, b) search engines. I promise this is the last, for at least, erm, 2 months. In my defence, I am simply highlighting an interesting article rather than making a bona fide blog posting!

Thursday, 22 October 2009

Blackboard on the shopping list: do Google need reining in?

Alex Spiers (Learning Innovation & Development, LJMU) alerted me via Twitter to rumours in the 'Internet playground' that Google is considering branching out into educational software. According to the article spreading the rumour, Google plans to fulfil its recent pledge to acquire one small company per month by purchasing Blackboard.

The area of educational software is not completely alien to Google. The Google Apps Education Edition (providing email, collaboration widgets, etc.) has been around for a while now (I think) and - as the article insinuates - moving deeper into educational software seems a natural progression and provides Google with clear access to a key demographic. This is all conjecture of course; but if Google acquired Blackboard I think I would suffer a schizophrenic episode. A part of me would think, "Great - Google will make Blackboard less clunky, offer more functionality and more flexiblity". But the other part (which is slightly bigger, I think) would feel extremely uncomfortable that Google is yet again moving into new areas, probably with the intention of dominating that area.

We forget how huge and pervasive Google is today. Google is everywhere and now reaches far beyond its dominant position in search into virtually every significant area of web and software development. If Google were Microsoft the US Government and the EU would be all over Google like a rash for pushing the boundaries of antitrust legislation and competition laws. This situation takes on a rather sinister tone when you consider the situation in HE if Blackboard becomes a Google subsidiary. Edge Hill University is one of several institutions which has elected to ditch fully integrated institutional email applications (e.g. MS Outlook, Thunderbird) in favour of Google Mail. Having a VLE maintained by Google therefore sets the alarm bells ringing. The key technological interactions for a 21st century student are as follows: email, web, VLE, library. Picture it - a student existence which would be entirely dependent upon one company and the directed advertising that goes with it: Google Mail, web (and their first port of call is likely to be Google, of course), GoogleBoard (the name of Blackboard if they decided to re-brand it!) and a massive digital library which Google is attempting to create and which would essentially create a de facto digital library monopoly.

I'm probably getting ahead of myself. The acquisition of Blackboard probably won't happen, and the digital library has encountered plenty of opposition, not least from Angela Merkel; but it does get me thinking that Google finally needs reining in. Even before this news broke I was starting to think that Google was turning into a Sesame Street-style Cookie Monster, devouring everything in sight. Their ubiquity can't possibly be healthy anymore, can it? Or am I being completely paranoid?

Monday, 19 October 2009

The Kindle according to Cellan-Jones

The world in which Rory Cellan-Jones exists is an interesting place. It's one which often results in a good, hard slap to the face. He can always be relied upon for some cynicism and negativity (or realism?) in his analysis of new technologies and tech related businesses. (See the last posting about Google Wave, for example.) This can be unexpected, often because he sees through the hype or aesthetics of many technologies and evaluates stuff based squarely on utilitarian principles. His overview of the Kindle is no exception to this rule:
"The Kindle looks to me like an attractive but expensive niche product, giving a few techie bibliophiles the chance to take more books on holiday without incurring excess baggage charges. But will it force thousands of bookshops to close and transform the economics of struggling newspapers? Don't bet on it."
The thing is, Cellan-Jones often talks a lot of sense. To be sure, the Kindle looks like an extremely smart piece of kit, but when Cellan-Jones stacks up the realities of the Kindle one wonders whether it'll be the game changer everyone is expecting it to be.

The focus for the Kindle seems to be on the best seller lists and the broad sheets. An area which appears to have eluded adequate exposition by all the tech commentators is the use of this new generation of e-book readers to deliver text books, learning materials, etc. This was always considered an important area for the early e-book readers. Why carry lots of heavy text books around when you could have them all on your Kindle or Sony Reader Touch, and be in a position to browse and search the content therein more effectively? Or, is this an extravagant use of E-Ink? E-Ink is required for lengthy reading sessions (i.e. novel) rather than dipping in and out of text books to complete academic tasks, something for which a netbook or mobile device might be better. So what happens to the future of e-book readers in academia?

Friday, 9 October 2009

Wave a washout?

This is just a brief posting to flag up a review of Google Wave on the BBC dot.life blog.

Google unveiled Wave at their Google I/O conference in late May 2009. The Wave development team presented a lengthy demonstration of what it can do and – given that it was probably a well rehearsed presentation and demo – Wave looked pretty impressive. It might be a little bit boring of me, but I was particularly impressed by the context sensitive spell checker ("Icland is an icland" – amazing!). Those of you that missed that demonstration can check it out in the video below. And try not to get annoyed at the sycophantic applause of their fellow Google developers...

Since then Wave has been hyped up by the technology press and even made mainstream news headlines at the BBC, Channel 4 News, etc. when it went on limited (invitation only) release last week. Dot.life has reviewed Wave and the verdict was not particularly positive. Surprisingly they (Rory Cellan-Jones, Stephen Fry, Bill Thompson and others) found it pretty difficult to use and pretty chaotic. I'm now anxious to try it out myself because I was convinced that it would be pretty amazing. Their review is funny and worth reading in full; but the main issues were noted as follows:
"Well, I'm not entirely sure that our attempt to use Google Wave to review Google Wave has been a stunning success. But I've learned a few lessons.

First of all, if you're using it to work together on a single document, then a strong leader (backed by a decent sub-editor, adds Fildes) has to take charge of the Wave, otherwise chaos ensues. And that's me - so like it or lump it, fellow Wavers.

Second, we saw a lot of bugs that still need fixing, and no very clear guide as to how to do so. For instance, there is an "upload files" option which will be vital for people wanting to work on a presentation or similar large document, but the button is greyed out and doesn't seem to work.

Third, if Wave is really going to revolutionise the way we communicate, it's going to have to be integrated with other tools like e-mail and social networks. I'd like to tell my fellow Wavers that we are nearly done and ready to roll with this review - but they're not online in Wave right now, so they can't hear me.

And finally, if such a determined - and organised - clutch of geeks and hacks struggle to turn their ripples and wavelets into one impressive giant roller, this revolution is going to struggle to capture the imagination of the masses."
My biggest concern about Wave was the important matter of critical mass, and this is something the dot.life review hints at too. A tool like Wave is only ever going to take off if large numbers of people buy into it; if your organisation suddenly dumps all existing communication and collaboration tools in favour of Wave. It's difficult to see that happening any time soon...

Thursday, 8 October 2009

AJAX content made discoverable...soon

I follow the Official Google Webmaster Central Blog. It can be an interesting read at times, but on other occasions it provides humdrum information on how best to optimise a website, or answers questions which most of us know the answers to already (e.g. recently we had, 'Does page metadata influence Google page rankings?'). However, the latest posting is one of the exceptions. Google have just announced that they are proposing a new standard to make AJAX-based websites indexable and, by extension, discoverable to users. Good ho!

The advent of Web 2.0 has brought about a huge increase in interactive websites and dynamic page content, much of which has been delivered using AJAX ('Asynchronous JavaScript and XML', not a popular household cleaner!). AJAX is great and furnished me with my iGoogle page years ago; but increasingly websites use it to deliver page content which might otherwise be delivered using static web pages in XHTML. This presents a big problem for search engines because AJAX is currently un-indexable (if this is a word!) and a lot of content is therefore invisible to all search engines. Indeed, the latest web design mantra has been "don't publish in AJAX if you want your website to be visible". (There are also accessibility and usability issues, but these are an aside for this posting...)

The Webmaster Blog summarises:
"While AJAX-based websites are popular with users, search engines traditionally are not able to access any of the content on them. The last time we checked, almost 70% of the websites we know about use JavaScript in some form or another. Of course, most of that JavaScript is not AJAX, but the better that search engines could crawl and index AJAX, the more that developers could add richer features to their websites and still show up in search engines."
Google's proposal involves shifting the responsibility of indexing the website to the administrator/webmaster of the website, whose responsibility it would be to set up a headless browser on the web server. (A headless browser is essentially a browser without a user interface; a piece of software that can access web documents but does not deliver them to human users). The headless browser would then be used to programmatically access the AJAX website on the server and provide an HTML 'snap shot' to search engines when they request it - which is a clever idea. The crux of Google's proposal is a suite of URL protocols. These would control when the search engine knows to request the headless browser information (i.e. HTML snapshot) and which URL to reveal to human users.

It's good that Google are taking the initiative; my only concern is that they start trying to re-write standards, as they have a little with RDFa. Their slides are below - enjoy!

Wednesday, 23 September 2009

Yahoo! is alive and kicking!

In a recent posting I discussed the partnership between Yahoo! and Microsoft and wondered whether this might bring an end to Yahoo!'s innovate information retrieval work. Yesterday the Official Yahoo! Search Blog announced big changes to Yahoo! Search. Many of the these changes have been discussed in previous postings here (e.g. Search Assist, Search Pad, Search Monkey, etc.); however, Yahoo! have updated their search and results interface to make better use of these tools. As they state:
"[These changes] deliver a dynamic, compelling, and integrated experience that better understands what you are looking for so you can get things done quickly on the Web."
To us it means better integration of user query formation tools, better use of structured data on the Web (e.g. RDF data, metadata, etc.) to provide improved results and results browsing, and improved filtering tools, something which is nicely explained in their grand tour. According to their blog though, better integration of these innovations involved a serious overhaul of the Yahoo! Search technical architecture to make it run faster.
"Now, here's the best part: Rather than building this new experience on top of our existing front-end technology, our talented engineering and design teams rebuilt much of the foundational markup/CSS/JavaScript for the SRP design and core functionality completely from scratch. This allowed us to get rid of old cruft and take advantage of quite a few new techniques and best practices, reducing core page weight and render complexity in the process."
I sound like a sales officer for Yahoo!, but these improvements are really very good indeed and have to be experienced first hand. It's good to see that the intellectual capital of Yahoo! has not disappeared, and fingers-crossed it never will. True - these updates were probably already in the pipeline months before the partnership with Microsoft; but it at least demonstrates to Microsoft why it still has the upper hand in Web search.