Showing posts with label user-generated content. Show all posts
Showing posts with label user-generated content. Show all posts

Monday, July 2, 2012

Twitter's war on partners: strength or weakness?

On Friday, LinkedIn announced it would no longer be distributing Twitter’s (140-character) content. Ryan Roslansky, “head of content products” helpfully pointed out that the policy only applied to content in, not content out: if you start a conversation on LinkedIn, it can be mirrored out to Twitter, but not vice versa.

Roslansky said the three-year collaboration was ended by Twitter, pointing to a blog post by a Twitter product manager entitled “Delivering a consistent Twitter experience” which was described as the nominal reason for the change to cut down on partner access to Twitter’s APIs and content.

Of course, this is brazen dissembling, as with similar claims last summer taking over URL shortening (bypassing bit.ly, tinyurl, etc) was to provide security. Twitter is kicking out partners because it wants control and to my mind that’s both a strength and a weakness.

The strength is that Twitter can do it — so far it’s the only game in town. If it aspires to end-to-end integration ala Google, the more partners it can disintermediate, the better. As with its website redesigns, this allows it to better know what’s going on — unlike the WWW, you can’t read anything on Twitter without signing in (and letting the company know whether you’re searching for Obamacare or Kardashian). If it has shut off LinkedIn, how far behind is Facebook?

At the same time, the increasing integration could also reflect a sign of weakness, specifically its one-trick business model. The only way Twitter makes money is putting eyeballs in front of ads. Any chance to view content away from Twitter is a chance to follow Twitter’s content without Twitter ads. That — and a natural fear of the 800 lb gorilla — is why it cut off Tweets from Google searches.

However, this increasing control has come at the expense of the user experience. I have long used the Tweetie client by Atebits LLC. Twitter bought Atebits, eliminated the client, and replaced it with a defeatured substitute posted to the Mac App Store. The website design also provides me less control of what I see and what I show on my website.

To my mind, this strategy is begging for competition. Twitter is gambling that nobody can knock it off the hill: the (direct) switching costs are low, but the network effects are huge. Google is certainly trying, but so far it’s not gaining traction and if Google doesn’t have the content (and customers) to displace Twitter, who does?

In the end, Twitter may become for me a write-only medium: I’ll give Twitter free content (that supports my blog and professional career) but not read content there. I suppose we both make out from the deal, but — like any other cranky old-timer — I’ll now and again remark how Twitter was once a useful news service back in the good ol’ days (when I walked 5 miles to grade school through the San Diego snow).

Saturday, November 28, 2009

We deserve better commodity information

It’s no news that Wikipedia, with all its flaws, is the default information source of a generation of skulls full of mush. If this wasn’t obvious enough from my college students, it was brought home a week ago when interviewing FLL robotics contestants (ages 9-14), when nearly all said their project “research” consisted of Google and Wikipedia. (One team said Google and Yahoo).

However, since then, Wikipedia’s problem has been a front page Wall Street Journal story Monday (and blog entry) on how Wikipedia is losing volunteers, specifically 49,000 in Q1 2009. The Telegraph had the most comprehensive follow up stories although the Times of London had good coverage (including a great article on the four sources of error.)

The impetus for the original WSJ article was the academic research of Felipe Ortega, who is part of a group studying open source software but actually did his Ph.D. dissertation on Wikipedia (a related but quite different species). He’s been tweeting to offer his comment on the current news coverage. While the bulk of his research hasn’t gone through the peer review, the abstract suggests he’s taken seriously all the research design issues.

After all the articles and the academic study, the official Wikipedia response is pretty unsatisfactory. It changes the subject, arguing that while the tide of new volunteers roughly matches the ongoing losses, at least the site traffic and number of articles continue to grow.

However, none of this relates to two inherent problems in Wikipedia that the current management is unable to solve, plus the third (and potentially catastrophic) outcome of WIkipedia’s commoditization of information.

The first problem is that Wikipedia publishes content by persistent idiots. Now that there are dozens or thousands of individuals trying edit articles on almost any topic, there are chronic edit wars with rival editors taking out each other’s changes in edit wars.

Competition is healthy — if there’s a selection mechanism based on quality or performance. Wikipedia has no such mechanism. Instead, what gets published comes from people who whine and bitch and moan, who win out over people who know what they’re talking about but have better things to do with their life. This works well for chronicling Simpsons episodes but not for summarizing academic research or major historical controversies. (Yes, I know that there are capable contributors, but in every battle between idiots and experts, the idiots are winning.)

I tried to sell this angle to a reporter I spoke with Monday, but I guess he thought it was just the griping of a snobby college professor who gave up years ago after watching his work be mangled by twits. However, in the 27 comments (thus far) to the official Wikipedia response were these five comments:

  1. Well, I have taken hours editing and polishing a biographical article about a scientist. There is nothing in the article now that is under dispute, yet it is probably going to be taken down and deleted as one editor is exercising his or hers petty power-plays.
  2. My most recent experiences have been quite negative: edits reverted with no reason, pages tagged as grammatically terrible when they were no such thing, or tagged as “not up to WP’s standards” when they were stubs and in some cases *still editing*. These taggings tended to be “drive-by” in the sense that some other editor dropped the tag onto the page or made their reversion but then failed to respond to explanations on the talk page for days.
  3. I used to spend a lot of time writing for Wikipedia, amending entries and creating new articles. Now it seems that a small number of self-appointed editors run the site. If I create new articles then they are nearly always deleted. If I correct information I know for a fact is wrong, it is reverted back and I am warned by the small sub class of elite editors
  4. I took on editing the Albigensian Crusade page a while back, a fairly simple job because what’s known about it comes principally from three contemporary chronicles dealing with the specific subject. A chronicle is self-indexed by time, therefore it should have been adequate to simply point readers in the direction of the sources, but no, that was inadequate, full references please. I got started, went so far, and checked if this was right. The %*$^^@# responsible refused to take the time to feedback, and was quite rude about it, so I stopped. Other appeals to administration went nowhere, and I concluded this is a system full of chiefs who can’t be bothered to get their hands dirty actually editing,
  5. I, for one, am one of those professional contributors who left Wiki in disgust. After spending a lot of time creating pages or adding a lot of content, some amateur came along and dumbed down the content and added fictious pictures that were purported to be of the creatures listed. It became a waste of my time to provide a lot of information that could be cut-and-paste into term papers, dissertations, reports, etc., and have some arm-chair contributor wreck it all.
This is a problem I’ve known about since soon after I joined Wikipedia in November 2003. The entire production process would have to be ripped up to fix this. Even Amazon has a way of providing feedback on user contributions so that readers know whose comments have been useful, even if it (and other processes) is fatally broken for highly polarized topics like politics.

One problem I didn’t see coming was the inevitable shift from original writing to maintenance mode. I started my main burst of Wikipedia contributions (2003-2004) by creating 11 new articles, from venture capitalists Eugene Kleiner and Tom Perkins to adding two missing campuses (CSULB, CSUSM) of the 23-campus CSU system.

Today, thanks to the law of large numbers (and the long tail) are very few significant articles left to be written. (Yes, Wikipedia has an article on only one Joel West — and it’s a lame one — but I don’t consider that a major omission.)

This reminds me of what I experienced in my first few years as a professional programmer: it is so much more more fun to write new code than maintain someone else’s code. In fact, as I became a manager I learned this is a major recruiting and staffing problem — even when you pay people, let alone when they’re volunteers. Over and over again, I saw that the manager or other “stuck” (high switching cost) programmers had to take the scut work so you could offer the new exciting stuff to attract the best talent.

Clearly, at Wikipedia existing volunteers don’t want to do the scut work, nor do the newcomers. If it’s de minimus, then (to use an analogy) perhaps good citizens will just pitch in and pick up the candy wrapper, but nobody’s going to spend a weekend clearing up the trash along the highway just for the fun of it.

Wikipedia is running out of good jobs to hand out. If you can’t give out fun work, how are you going to attract people? What I didn’t see six years ago was that inevitably Wikipedia’s content base would mature: first in English and eventually in all the major languages. When this happened, the opportunities for adding new content would mainly be limited to current events like new hurricanes or those Simpsons episodes.

However, I find hope in Wikipedia’s current troubles, as they suggest a solution WIkipedia’s most invidious problem: the commoditization of human knowledge. Monopolies are bad, even if they are for free goods. When I was interviewing open source leaders, the Apache (and most “open source” types) seemed to get this, while the free software types (Linux, OpenOffice) did not.

Competition is inefficient, but it provides choice. Monopolies at best mean benevolent dictators, and few benevolent dictators remain benevolent forever.

The mind-numbing ubiquity of WIkipedia is teaching a generation of kids to be lazy and uncritical consumers of information — whether it’s truth or merely wikitruth. They take what shows up on the first page of Google or in Wikipedia and assumes it’s true, even when it’s not.

When I was a kid, I would do my 5th grade reports using World Book, Encyclopedia Brittanica, usually one other encyclopedia like Collier’s or Compton’s, and also the Information Please Almanac. (If the report was important, I would also try to find a real book or two.) This wouldn’t make me an expert, but at least I would get multiple perspectives.

Today, Wikipedia’s commoditization of information means that Encyclopedia Britannica is struggling and its previous nemesis (the Encarta CD-ROM) is gone. At least a five-year-old version of the Columbia Encyclopedia survives as Reference.com.

Once upon a time, I assumed that the network effects meant that nothing would ever compete with Wikipedia. This week shows that in less than a decade it’s possible to create a significant body of knowledge with volunteer labor. None of the existing rivals have yet succeeded, whether Citizendium, Conservapedia, Liberapedia or Knol. However, with this large body of existing (or potential) body of would-be Wikipedia labor becoming available, they are certainly trying.

I will be curious to see if we can achieve success from volunteer organizations that focus on the quality rather than the quantity of contributions. In this direction, Citizendium (by WIkipedia co-founder Larry Sanger) is using a somewhat modified version of the Wikipedia process, while Google’s Knol is heading in a different direction by emphasizing authorial integrity over cumulative production.

Given the almost total lack of competition, anything that provides a viable alternative to WIkipedia is a good thing. It will be a good thing if a decade from now we have three or four online encyclopedias to choose from, much as today we can choose from three or four cellphone carriers.

It’s likely that one of these alternatives will be Wikipedia. Perhaps if its leaders take its current problems seriously, it will still be the most popular alternative out there and will be able to meet its current modest fundraising goals.

Friday, February 6, 2009

GoDaddy on top?

What was the most effective SuperBowl ad? It depends on how you measure it.

USA Today used focus groups and concluded that it was the $2,000 “snow globe” Doritos ad, which had won its “Crash the SuperBowl” content for user generated content — in this case user-submitted advertisements. Joe and Dave Herbert from Ohio who made the ad collected $1 million for their efforts, and were on the Tonight Show earlier this week. This is also the most popular ad on YouTube.

2nd and 3rd on the USA Today list went to two Clydesdale ads, 4th to Mr. Potato Head — all ads I thought were effective. Another Doritos spot placed 5th, while three of the next four were ads I thought relatively effective: Cars.com overachiever, Pepsi’s “Forever Young,” and the Castrol grease monkeys.

Citing increases in gameday web traffic, compete.com said Denny’s was far and away the winner, with Cheetos, Pepsi, Bud and Gatorade (the big G?) far behind. ComputerWorld tried to estimate the most popular online ads, and their top three were the CareerBuilder ad (which I though crushed the Monster ad), the snow globe and the Conan O’Brian ad that I said was “funnier than his late night show usually is.”

TiVo measured the most watched ads on its time-shifting boxes, and came up with a slightly different list. The Doritos snow globe ad fell to fourth, and none of the other USA Today top 10 made the list. Number one was “Enhanced,” the slightly more effective and less offensive GoDaddy ad that was buried near the end of the game.

However, rewound ads don’t necessarily mean effective ads. Perhaps they are a measure of attention, but (as TechCrunch argued a year ago), they could easily be a measure of an ad that people didn’t understand. Which means they were of below-average effectiveness for non-TiVo customers.

At least TiVo concluded that the ETrade baby was only a one-year wonder. Since he’s outgrown his usefulness, let’s hope his parents will send the brat to preschool to get socialized.

Friday, January 23, 2009

Aggregating UGC

On the WSJ, Carl Bialik writes a very interesting column (and blog) called The Numbers Guy, both of which talk about the use and misuse of numbers to convey information (or psuedo-information).

This morning’s column (inside the paywall) talks about the controversy within the movie critic profession over the use of “star” ratings, and he elaborates on it (outside the paywall) in his blog. In addition to the impact on user generated content (more later), the column caught my attention as the former Arts Dept. editor for The Tech, the MIT student paper.

Both address the issue of reducing everything to one number. At The Tech, we used turkeys instead of stars, reverse coded from 0-5 (no turkeys is best). Other systems use 1-4 or 1-5 stars. So if a movie wins 3/5 stars, is it an average movie across the board, or is it a movie that’s great on some aspects (e.g. performances) and terrible on others (script)?

So this problem is one of “good” being a multi-dimensional measure. Hotels.com and other consumer rating sites seem to be able to compile these multiple dimensions on their UGC and allow buyers to see ratings on the dimensions that matter to them.

Hotels.com, Consumer Reports and others also have the issue of norming for price. Does 4 stars mean the same for a $50 motel and a $500 hotel? (Or for a $15,000 or $50,000 car). For the AAA diamond system, it’s an absolute scale, but Hotels.com seems to want the rating normed based on value.

As Bialik notes in his column (and an earlier column) the problem at the next level of analysis: aggregating the ratings of multiple reviewers, as is done by Rotten Tomatoes and Metacritic. Does a 3/5 mean everyone gave it a 3, a bell-curve distribution around 3, or a bimodal distribution of all 1 or all 5?

This is nicely (& easily) handled by Amazon, which gives both a mean and a histogram. This works particularly well for polarizing authors like Al Franken or Ann Coulter. However, it doesn’t fix the sampling bias issue — the people who self-select to write a review are not representative of the reading public overall (an issue Bailik notes in his paid article).

One issue that Bialik doesn’t address is weighting the average. For movies, critics who see dozens or hundreds of movies a year usually reward things that “push the envelope,” which normally means some sort of aberrant production values (Blair Witch), script (Curious Case of Benjamin Button) or characters (Boys Don’t Cry). Others are excited by the craft — the production values, acting performance, directing — of the sort that win Oscar® awards.

However, I am plunking down $10 2-4 times a year to be entertained. I want something that I enjoy watching, not something that pushes the envelope. A good example was “You’ve Got Mail,” which is a very nicely done romantic comedy — the ideal date movie — but only got a 62% critic rating. Sure, the plot was predictable, and the ending was pre-ordained, but the character quirks and plot twists were believable and at times funny: not as timeless as Tracey-Hepburn, but as close as a modern-day rendition as Hollywood offers nowadays.

So even if we have an accurate, low variance attribute rating — everyone agrees this is a predictable plot — it doesn’t solve the problem that buyers differ on how important that is. The problem was solved 30+ year ago by Fishbein & Ajzen, who described a subjective utility model with different attribute importance ratings (i.e. weights).

To discern what’s important to each buyer, this seems like a job for a neural network or other self-training system. Again, the Amazon recommendation engine handles this: if I like Al Franken, it shows me Bill Maher and Stephen Colbert; if I like Ann Coulter, it shows me Rush Limbaugh and Bill O’Reilly.

Interestingly, Amazon has very different incentives than the movie sites in aggregating ratings. Rotten Tomatoes or Metacritic want me to linger on the site looking at ads. Amazon wants me to buy something that I like, so I’ll buy more. Whether it’s different incentives or a different scale of revenues, Amazon seems to be much more serious about giving recommendations that are accurate for my tastes.

Perhaps what we need is an open source score aggregation system, one that handles

  • multiple dimensions or rating
  • conveys the distribution of ratings, not just the average
  • fits my own opinions to those of reviewers whose opinions most closely match mine, to give me feedback consistent with my tastes, interests and values.
Slightly off-topic: oddly for a numbers guy, Bailik uses the word “commodify” when he means “commoditize”. Google ranks the latter as 4x as popular. Wikipedia (not the most reliable source) attributes commodify to Marxist political theory, while accurately noting that commoditize is the term used in business.

Tuesday, December 16, 2008

Users lie

When we had a tech support operation, our former product dictator trained the tech support people with the maxim: “users lie.” I thought it was a bit extreme, but then this manager liked extreme rhetoric for effect.

The underlying truth is that when handling a tech support calls, the information from users is not 100% reliable. (As someone now on the other side again, I can certainly see that.) Users may leave something out and tell the story slightly inaccurately. Or they might deliberately leave out relevant information:

Q: When did your cell phone stop working?
A: I dunno, it just won’t turn on anymore.
instead of
Q: When did your cell phone stop working?
A: When I dropped it on the concrete sidewalk.
(I would have used the example of dropping in water, but cellphone companies now test for that.)

Michael Mace highlights another, curious example of deliberate misrepresentation — political retaliation over Proposition 8 (the initiative repealing the June court decision legalizing gay marriage). In this case, users can lie when generating user-generated content.

The why is pretty simple: Opponents of Prop 8 want to retaliate against any individual or business that supported Proposition 8. The example he uses is a Mexican restaurant where the manager gave $100 to Prop 8 and then gay marriage advocates retaliated on Yelp. (Of course, this might extend to other social or political controversies, such as a non-union grocvery store)

It seems to me there are at least three cases:
  1. Customers who complain about a business but then give ta numerical rateing that the store aotherwise deserved. While iritating, this does proviade additional information that might be important to some customers. The risk, of course, is if the sample of activist-complainers is biased: pro gay marirage types provide information but pro life (anti abortion) types do not.
  2. Customers who complain about the busienss polciies and then give a dishonest rating of the quality of the firm’s goods or services in a deliberate attempt to lower the numeric scores. Mike shows a lot of this going on in this case
  3. Customers who completely lie; the hyperbole of this example that Mike quotes is emblematic:
  4. “I’ve been here a few times, and this is without a doubt the worst place I know of in California. Do not go here unless you don’t mind bad food, high prices, a horrible vibe, and can turn your back on risks of food poisoning and human rights and health code violations!”
The more you try to stamp out #1 and #2 (which are recognizable), the more you get #3 (which are harder if not impossible to recognize.

Mike writes not about the implications for restaurants, but for the reliability of user rating sites. The problem is, such sites were never all that reliable in the first place. For example, I went searching for a hotel for a possible (now unlikely) Hawai‘i trip, and found a few places where the only rating was an anonymous 5-star with no comments; presumably these are posted by the owners or their friends. In fact, hotel sites in general seem to have unreliable ratings.

Yelp could go for weighting the contributions of users (useful/not useful) as does Amazon, but then this could becoming an ideological war of fellow travelers vs. enemies. Look at the reviews of any politically or socially controversial book on Amazon — any book (or comment) that takes a stand will become a battleground for ideologies rather than the merits of the book.

Some e-commerce sites will only allow you to rate what you’ve bought (and they can verify that). Although not applicable to independent ratings site, that does reduce abuse. However, the reduction of quantity means little or no feedback on thinly-traded goods and services.

In the end, this is exactly the same problem as Wikipedia or any other site built upon user-generated content. The site designers generally assume altruistic self-disclosure, and fall down (if not fail miserably) at any systematic source of bias in the evaluations. The goal of encouraging maximal coverage through maximal participation is in direct conflict with having control over the quality. The problem is magnified by fragmentation of contributions across many sites, thus increasing the pressure to attract contributions (of unknown quality) at all costs.

The incentives (and energy) to lie are stronger than those to catch the lie. Manual catching systems don’t scale; attempts to use peers to discover bias will only work if the peers also are free from bias (they aren’t). Algorithms might work for a while, but those who game the system will discover ways to make their bias more difficult to detect.

The problem will only get worse: as crowd sourcing and its impact becomes better understood, efforts to manipulate it become better organized, more common and probably more subtle (and thus hard to detect). It’s possible (by no means assured) that a decade from now crowd sourcing will prove to be a noble experiment that failed.

I think the end result will be to bring us back full circle, to where we were a decade ago: I will trust only the feedback of someone I know, either personally or a brand name reviewer or analyst who has proven reliable (by my standards) in the past. I’m not sure how that helps me find a restaurant in Los Angeles or Scottsdale, although my standards for under $10 restaurants are pretty lax.