Be a Top Mechanical Turk Worker: You Need $5 and 5 Minutes

The current reputation system on Mechanical Turk is simply inadequate. The only built-in reputation metrics are the number of completed HITs and the approval rate. 

Some people believe that they are adequate as a basic filtering mechanism. They are not. 

For example, ask for all workers in your HITs to have 1000 completed HITs and 99% approval rate. You believe that you will only get high quality workers? You are wrong!

I tried to filter workers using just these metrics. I failed. Spammers got me again. (And once in, spammers submit a lot of crap. It costs nothing.) I questioned why. How can it be? And then I realized: It is trivial to beat these metrics.

Let's see the effort it takes to beat the system. 

The mission: Become a top Turker, 100% approval rate and 1000 completed HITs. 
  • Step 1: Login as a requester. Post a task, with 1000 HITs. Each HIT pays 1 cent. Total cost: $15. Out of these, $10 go to the worker, $5 go to Amazon. The title of the HIT: "Write a 500 word review". No sane worker will touch these HITs. Done. Logout.
  • Step 2: Login as worker, using a different email. Complete and submit the 1000 HITs created in Step 1. You just need to click submit 1000 times. Bored? iMacro and Greasemonkey can help. Done. Logout.
  • Step 3: Login as a requester again. Approve all submitted HITs. Pay the $15. Amazon gets $5. The worker account has the remaining $10. Done. Logout.

Your worker account is a top Turker now. 1000 completed HITs and 100% approval rate. Congratulations! You have a license to spam.

The Explosion of Micro-Crowdsourcing Services

In my last post, I expressed my surprise for the sudden explosion of the research-oriented workshops in computer science conferences that are explicitly focused on the concept of crowdsourcing.

I should also note though, that there is a parallel explosion of similar micro-crowdsourcing services. Here is a list of services that I have ran into: 
Some of the companies above are serious, some are new and upcoming, some are copycats, and some are there just to facilitate spamming. 

I thought of doing a more detailed comparison (similar to the report that Brent Frei prepared last year for the more general area of paid crowdsourcing) but then I realized that I do not trust enough half of these companies to even give them my email.

This growing list makes it clear that we enter the bubble period. Bubbles are not necessarily negative. During bubble periods you see many innovations coming into the industry from many different parties. While most of the entrants in the market will die sooner rather than later, I except to see interesting things coming out of this. 

Do not forget that the dotcom bubble generated the Pets.com failures but also gave birth to Google, who replaced the early dominant players, such as Lycos and Altavista. 


The Explosion of Crowdsourcing Workshops

Over the last couple of years, there has been an explosion of workshops that look at the topic of crowdsourcing from the academic point of view, within the broader computer science field. Here are the ones that I am aware of:
  1. Human Computation Workshop (HCOMP 2009), with KDD 2009
  2. Workshop on Crowdsourcing for Search Evaluation, with SIGIR 2010
  3. Second Human Computation Workshop (HCOMP 2010), with KDD 2010
  4. Advancing Computer Vision with Humans in the Loop (ACVHL), with CVPR 2010
  5. Creating Speech and Language Data With Amazon�s Mechanical Turk, with NAACL 2010
  6. Computational Social Science and the Wisdom of Crowds, with NIPS 2010
  7. The People�s Web Meets NLP: Collaboratively Constructed Semantic Resources, with COLING 2010
  8. Workshop on Ubiquitous Crowdsourcing, with UBIComp 2010
  9. Enterprise Crowdsourcing Workshop, with ICWE 2010
  10. Collaborative Translation: technology, crowdsourcing, and the translator perspective, with AMTA 2010
  11. Workshop on Crowdsourcing and Translation
  12. Crowdsourcing for Search and Data Mining, with WSDM 2011
(If you think that I missed a relevant workshop, drop me a line, and I will add it to the list above)

In addition to the workshops above, we also have the CrowdConf 2010 conference, organized by CrowdFlower, with some academic presence but overall targeted mainly to industry.

Yes, one workshop in 2009, followed by ten (at least!) additional workshops in 2010, and who knows how many more in 2011. (I am already aware of 3 planned workshops, in addition to the one in WSDM 2011.) 

I am deeply interested in the topic and I already feel that I am losing track of the venues that I need to follow.

Mechanical Turk Requester Activity: The Insignificance of the Long Tail

The Pareto principle says that 80% of the effects come from 20% of the causes. It is a favorite anecdote to cite that 20% of the employees in an organization do 80% of the work, or that 20% of the customers are those that generate 80% of the profits.

In online settings, such inequalities are often amplified. For Wikipedia we have the 1% rule, where 1% of the contributors (this is 0.003% of the users) contribute two thirds of the content. In the Causes application on Facebook, there are 25 million users, but only 1% of them contribute a donation.

So, adapting this question for Mechanical Turk, we want to see: What is the distribution of activity across requesters?

The Activity Distribution: The (Insignificance of the) Long Tail of Requesters

To analyze the level of participation, for the XRDS paper, we took the requesters that posted a task on Mechanical Turk from January 2009 until April 2010, and we ranked them according to the total reward amount of the posted HITs. Then, we measured what percentage of the rewards comes from the top the requesters in the market. Here is the resulting plot:


Indeed, the result shows that Mechanical Turk is closer to the "1% rule" of Wikipedia, than to the general 80-20 principle. As in Wikipedia, the top 1% of the requesters, contribute two thirds of the activity in the market.

By reading the graph, we see the following:
  • Castingwords, the top requester across the 10K requesters in the dataset, accounts for 10% of the dollar-weighted activity (!).
  • The top 0.1% of the requesters (i.e., the top-10 requesters) account for 30% of the dollar-weighted activity.
  • The top 1% of the requesters account for 60% of the dollar-weighted activity.
  • The top 10% of the requesters account for 90% of the dollar-weighted activity.
  • The long tail of the 90% of the requesters is effectively insignificant.
A closer look at the distribution of requester activity shows that the activity per requester follows roughly a log-normal distribution.

  • The average level of posted rewards is $58. This corresponds to an average level of activity of just four dollars per month.
  • The median is just $1.60. Yes, this is not a typo: 1.6 dollars. In other words, 50% of the requesters never post more than a couple of dollars worth of tasks.
  • Only a small fraction of requesters (less than 1%) posted 1000 dollars worth of tasks or more over the period from January 2009 till April 2010. 
The lognormal distribution of activity, also shows that requesters increase their participation exponentially over time: They post a few tasks, they get the results. If the results are good, they increase by a percentage the size of the tasks that they post next time. This multiplicative behavior is the basic process that generates the lognormal distribution of activity.

I would like to try is to check if this model indeed corresponds to reality. Do we see a geometric growth in activity as the requester stays in the market for longer? Do we observe "deaths" of requesters? (The Fader-Hardie model may be a nice, simple model to try.) What is the expected future activity of a requester?

Such questions may be useful for guiding decisions of workers when deciding whether to invest time and effort to get a good reputation for a given requester (e.g., by completing qualification tests or completing the basic HITs that unlock access to the "protected" HITs.)

What Tasks Are Posted on Mechanical Turk?

A few months back, I got an invitation from Michael Bernstein (of Soylent fame) to write a small article about Mechanical Turk for the student magazine of ACM, the ACM XRDS (aka Crossroads). I could have written a summary of past research, a position paper, or anything that I find interesting.

Instead of summarizing and resubmitting already published material, I decided to push myself and start analyzing some data that I have been collecting about the Mechanical Turk marketplace. In the past, I analyzed the data about the demographics of the workers on Mechanical Turk. However, we have limited analysis for the requester side of the market, and for the type of tasks being posted.

My goal was to put some very preliminary analysis in place, just to scratch the surface of a variety of a questions that I have heard over time. Hopefully this will push me to start working more towards getting the answers, and will inspire some interesting new questions for the students and others that read the article. A preprint of the paper is available through the NYU Faculty Digital Archive and the print version should appear sometime early in 2011.

Data Set

I have been collecting data about the marketplace through my Mechanical Turk Tracker. The tracker collects complete snapshots of the marketplace, every hour, starting from January 2009. For the analysis, I took the data from the period of January 2009 till April 2010: The snapshots has a total of:
  • 165,368 HIT groups
  • 6,701,406 HITs
  • 9,436 requesters
  • $529,259 rewards
These numbers, of course, do not account for the redundancy of the posted HITs, or for HITs that were posted and disappeared between our hourly crawls. Nevertheless, they should be good approximations (within an order of magnitude) of the activity of the marketplace.

Top Requesters

The first question that I looked at was an analysis of the tasks that are being posted on Mechanical Turk. One way to understand what types of tasks are being completed in the marketplace is to find the �top� requesters and analyze the HITs that they post. By ranking requesters according to the sum of the posted rewards, we get the following list, showing the level of activity and the type of tasks that these requesters post. (Note: To avoid skewing the data towards one-shot requesters, I excluded from the list a requesters that were active only for small periods of time or requesters that posted only a small number of HITs. The goal was to find not only the requesters that post big tasks, but also requesters that do that consistently over time.)


So, transcription, classification, and content generation seems to be a common activity on Mechanical Turk. This indicates that people have developed sufficient best practices and can actually get quality work done. (If not, they would not be posting so many tasks.)

Top Keywords

We also wanted to get a feeling of the tasks that are being posted in the market, across all requesters. The table below shows the top-50 most frequent HIT keywords in the dataset, ranked by total reward amount, # of HITgroups, and # of HITs.


Beyond the tasks identified before, we also see data collection, image tagging, website feedback, and usability tests to be common tasks being posted in the marketplace.

In future posts, I will post further analysis of other aspects of the AMT marketplace.

 
Free Host | lasik surgery new york