The New Demographics of Mechanical Turk

Past surveys on the demographics on Mechanical Turk users indicated that most of the workers come from the US, are younger and more educated than the general population, and work on MTurk as a way to get some spare cash.

Since the last survey, a few things have changed. First, Amazon allows now workers in India to get paid in cash in rupees, essentially encouraging many people from India to start using Mechanical Turk as workers. Second, the recession has affected many households, leaving many people at home looking for cash to cover their needs. These two forces has changed the demographics of the participants, so a new survey was needed to capture the new demographics of the Mechanical Turk workers.

So, in February 2010, I conducted a new survey on Mechanical Turk, paying the workers10 cents for participating.

The first major change was the country of origin. In the past 70%-80% of the workers were coming from the US, but now the percentage is closer to 50% and it may decrease even more. India is now a major contributor of workers, with almost 35% of the workers coming from the subcontinent. The remaining workers come from 66 different countries. The exact numbers in the survey:


  • United States: 46.80%
  • India: 34.00%
  • Miscellaneous: 19.20%

The analysis also indicated that the profile of the Indian workers is quite different from the profile of the U.S-based workers. So, below I present the results broken down by country.

Gender Breakdown

The first analysis focus on the gender breakdown. Across US-based workers, there are significantly more females than males, while the situation is reversed for Indian workers.



The main reason for the overrepresentation of females in the US-based workforce is the nature of the tasks and work on Mechanical Turk. Most participants in the US use Mechanical Turk as a supplementary source of income, and often Mechanical Turk is used by stay-at-home parents, unemployed and underemployed workers, and so on. Since females are more likely to fit into these categories, there is a corresponding increase in representation. On the contrary, more Indian workers treat Mechanical Turk as a primary (or at least significant) source of income, and we see more males working on Mechanical Turk.


Age Distribution

In terms of age distribution, there is definitely an overrepresentation of younger workers, compared to the general population of Internet users. While this holds both for the US and for India, we see an even higher skew towards younger workers among Indians.






Educational Level

We also asked the Mechanical Turk workers to declare their educational level. In general, the (self-declared) educational level of the workers is higher than the general US and Indian population. There are two factors that may contribute to this. First, many of the workers are younger than the overall population and, ceteris paribus, this leads to higher educational level. Finally, while we may not necessarily discount the possibility of false disclosure, there are no incentives that would bias workers towards lying in this survey.




Income Level

We were also interested to examine the income level of the workers on Mechanical Turk. In the US, the shape of the distribution roughly matches the income distribution in the general US population. However, it is noticeable that the income level of US workers on Mechanical Turk is shifted towards lower income levels. For example, while 45% of the US Internet population has income below $60K/yr, the corresponding percentage across US-based Mechanical Turk workers is 66.7%. (This finding is consistent with the earlier surveys that compared income levels on MTurk workers with income level of the general US population of Internet users.) The picture is drastically different across US-based and Indian workers. Workers based in India have significantly lower incomes, as expected, and more than 55% of the workers declared an income of less than $10K/year.




Marital Status, Children, and Household Size

In terms of marital status and household size, the answers tend to match the age demographic of the workers reported earlier. The majority of the workers, both in India and in the US, do not have children, and a significant fraction of them are single. An interesting contrast is the household size, which seems more to reflect cultural norms than anything specific to Mechanical Turk: While more Indian workers are single and without children, they seem to stay in houses with larger number of household members, compared to US workers: Indian workers either stay with their family, or they tend to have a comparatively larger number of roommates, compared to US workers.




Level of Engagement on Mechanical Turk

We also asked a set of questions for evaluating the level of engagement of Mechanical Turk workers on the marketplace. Since we did not detect significant deviations across countries, we will be reporting the results in aggregate form, without separating by country of origin of the worker. In general most workers spend a day or less per week working on Mechanical Turk, and tend to complete 20-100 HITs per week. Correspondingly, this generates a relatively low income stream for Mechanical Turk work, which is often less than $20 per week. Of course, there are a few workers that devote a significant amount of time and effort, completing thousands of HITs, and generating a respectable income of more than $1000/month. For these workers, Mechanical Turk tends to be the primary source of income, of course. For Indian-based workers, such salary levels are typically satisfactory for the type of work that is available on Mechanical Turk (i.e., tedious tasks that do not require significant specialized skills)




Motivations for Participating on Mechanical Turk

To understand better why people participate on Mechanical Turk, we asked for both qualitative (i.e., free text) and a set of structured questions. The main structured question that we asked was the following:

Why do you complete tasks in Mechanical Turk? Please check any of the following that applies:
  • Fruitful way to spend free time and get some cash (e.g., instead of watching TV)
  • For "primary" income purposes (e.g., gas, bills, groceries, credit cards)
  • For "secondary" income purposes, pocket change (for hobbies, gadgets, going out)
  • To kill time
  • I find the tasks to be fun
  • I am currently unemployed, or have only a part time job

The answers were quite different across Indian and US-workers. Very few Indian workers participate on MTurk for "killing time", and significantly more Indians treat MTurk as a primary source of income. (Not surprising given the average income level of an Indian worker vs the income level of the US workers.)




While these graphs are ok, I would actually encourage everyone to go through the textual responses of the workers. Below you can find the data embedded in a Google Spreadsheet. Go through the column "EngagementQ1" and I am sure that you will enjoy reading all the answers given by the workers.



More Details

The set of blog posts about the demographics of Mechanical Turk have been a little bit too popular, reaching a point where people were asking me how to cite these surveys. While I thought that pointing to the blog would be enough, I was surprised to find out that many people consider a blog post to be too "informal" to cite.

Although I find this academic conservatism kind of funny, I am not sure whether such type of work should appear in an academic conference, journal, or magazine. Maybe a magazine-style journal would be ok but I am still not 100% convinced.

Anyway, as a compromise, I now created a working paper with the results of these demographic studies, and now whomever is interested can cite the "official" working paper with the results on the demographics of Mechanical Turk. (As you will notice, it is essentially this blog post, pasted in a PDF file.) As an added value, you can also find there the Excel spreadsheet with all the results and perform your own analyses and studies.

I would like to consider this working paper as a true working paper, i.e., update it over time with the most current results, if I consider this necessary. Until then, enjoy the current results and let me know if there are other questions that you would like to see answered.

Why Mechanical Turk Allows Only US-based Requesters?

Many people read the blog from outside the United States. All these readers learn about Mechanical Turk and are excited about the concept, so they want to try it out. Unfortunately, no such luck! You need a US credit card to be able to fund the account (or you need to work as a Turker to accumulate the amount necessary to fund your own tasks.)

So, many people are asking: Why Amazon does not open the service internationally? Why restrict Mechanical Turk only to people that have US credit cards?

This was kind of puzzling to me as well, given that other Amazon Web Services (e.g., EC2, S3, etc) are open to international customers. Why other web services are open but MTurk is not?

Today I met with John Hoskins, Senior Manager of Business Development of Mechanical Turk, and asked the very same question. The answer was clarifying: For all other web services, the customer is consuming Amazon services and pays Amazon. For Mechanical Turk, Amazon receives funds from requesters and then distributes them to workers.

This flow of payments forces Amazon to comply with the US Patriot Act, especially the provisions about money laundering and financing of terrorist activities. The basic idea, known as the "Know Your Customer (KYC)" doctrine, is that Amazon should know from whom they get money and to whom they send the money. This is possible for US credit cards and for US bank accounts, due to the regulations of the US banking system. (I also guess that this is also possible for India, given that Amazon now pays workers in India using rupees.)

I find it kind of fascinating that Mechanical Turk could be used as a venue for money laundering but, in retrospect, not unlikely. In fact, given the relatively low fees that Amazon charges for funds to change hands, it is almost appealing. I can easily see a person posting one million $10 do-nothing HITs, available only to a single qualified worker, who can then consume them and get paid "clean" money.

I guess the news is a disappointment for many aspiring international requesters, but there is some hope: If you can open a US-based credit card (e.g., have a pre-paid credit card or a gift credit card), then it should be possible to open a Mechanical Turk requester account. Or, simply go and use CrowdFlower that will serve as an intermediary and submit your tasks to Mechanical Turk, providing many other value-added services along the way (thanks to S�rgio Nunes for reminding me about that!).

Universities and Intellectual Property: A Minefield?

One of the things that I never understood at NYU is what are the rights that the university has on the work produced by the faculty members and students.

Following the intellectual properly law, there are four basic types of intellectual properly:
  • Copyright
  • Patents
  • Trademarks
  • Trade Secrets
If someone works for a corporation, things are pretty clear. Any paper, before being published needs to get approval. Any developed algorithms and code written is the intellectual property of the company and the company owns the copyright for the code, and can treat the algorithms as a trade secret. The company may also patent useful inventions and register some valuable trademarks. But, in all these cases, everything that is being produced within the corporation is work made for hire and owned by the corporation. The employee has typically no ownership of the produced work and it is commonly prohibited for the employee to work for another company or provide any sort of consulting services..

For the work of faculty, I always felt that everything falls into a grey area. Most of the work is made public as soon as possible. Code is often released as open source, following some pretty liberal licensing scheme, or even released to the public domain. Papers are written and publicized without much, if any, vetting and the algorithms and methods described there are typically in the public domain. The only case where a university has some control over intellectual property is when a patent is filed and granted.

Now, the great confusion arises when the faculty wants to work with a corporation and the university allows faculty members to engage into consulting agreements. Who owns and controls the expertise and discoveries of the faculty member?

Let's say that myself, Panos, invented an algorithm in area X, wrote a paper, and published the code in an open source format. Corporation A, comes to me and asks me to consult them on area X. What is the control that my employer, NYU, has on my work? Yes, Corporation A wants to hire me because of the IP that I produced while at NYU. This IP though is publicly available, so I do not really transfer anything protected under copyright law.

I have asked this question to our own tech transfer office. Unfortunately, I did not get back a clear answer. They told me that I cannot transfer code and that any patent is owned by NYU. Correct, these are indeed intellectual property assets. (Although for the case of open source code, this is again confusing.) But what about the expertise that a faculty member develops? In corporations this is often protected using some no-compete clauses in the employment contract, effectively preventing employees from directly transferring know-how. In universities, there is no such provision.

I find this merging of academic and corporate worlds to be particularly confusing and I find this to be a potential minefield. Who owns what? Any ideas? Any experiences? How other universities treat the concept of tech transfer?

Did you find this helpful?

Last week, the New York Times Sunday magazine had an article titled The Reviewing Stand, starting with the following:

Here�s a challenge for students of expository writing: review a popular product on Amazon and aim to get your review chosen by readers as �most helpful.� It�s dead hard. The product review, as a literary form, is in its heyday. Polemical, evocative, witty, narrative, exhortative, furious, ironic, off the cuff....

What I found amusing was the fact that, after reading this article, I got a notification that the journal version of the paper Estimating the Helpfulness and Economic Impact of Product Reviews: Mining Text and Reviewer Characteristics, co-authored with my frequent co-author, Anindya Ghose, has been accepted for publication at the IEEE Transactions on Knowledge and Data Engineering (TKDE) journal.

As the title suggests, one of the problems that we attack in the paper is how to predict the usefulness of a product review. For example, if you go on Amazon, you will see, on top of many reviews, how many people considered a particular product review helpful:



So, the question is: Can we predict how helpful a particular review will be?

Our first attempts to address this problem appeared in the WITS 2006 and the ICEC 2007 papers. Following the scientific zeitgeist, a large number of other papers appeared these years, all tackling the question of predicting helpfulness of reviews. (See the actual paper for references.)

What I found rather surprising was the relative easiness of the task. A few relatively straightforward features can be used to predict with good accuracy whether a review will be deemed helpful or not.
  • Check the readability of the article, as measured by one of the many readability metrics, check the number of spelling errors, and measure basic statistics of the text, such as review length. Using just the readability and the fraction of spelling errors in the article we can estimate with 70%-80% accuracy whether a review will be deemed helpful or not.
  • Check the history of the reviewer. If the reviewer has been writing helpful reviews in the past, it is highly likely that reviews in the future will also be helpful. Also, if a reviewer has disclosed personal details (name, location, etc) the reviews are more likely to be helpful. Again, using just reviewer history and disclosure details, we get 70%-80% accuracy, as measured with the AUC metric.
  • Check the "subjectivity" of the review. We call a review objective if it contains mainly information that can be found in the product description and specs. A subjective review contains information that depends on the personal experiences of the reviewer. Helpful reviews tend to contain a mix of both.
Interestingly enough, all three feature sets seem to have equivalent predictive power. Even using them all together does not seem to increase substantially the predictive performance.

While preparing the final version of the paper, I also checked other papers that were attacking the same problem. While many papers were trying to predict helpfulness using textual features, I noticed that a few papers were using a set of alternative and interesting features:
  • Coverage of product features. Many products can be considered an aggregation of multiple product features. For example, a digital camera has resolution, size, battery life, sensor size, etd. How many product features are being discussed in the review? This feature tends to have predictive power, according to (Liu et al, EMNLP 2007).
  • Dynamics of reviews. Reviews that are posted early on get a higher fraction of helpful votes. In contrast, later reviews need to be more informative and comprehensive to attract the same fraction of helpful votes (Liu et al, EMNLP 2007). 
  • Controversy. The helpfulness of a review depends not only on its own content but also on how controversial is the product under consideration (Danescu-Niculescu-Mizil, WWW 2009).
  • Social network of reviewers. If reviewer A trusts the reviews of reviewer B, then the reviews of B are likely to be more helpful than the reviews of A. ("Exploiting Social Context for Review Quality Prediction"; by Lu, Tsaparas, Ntoulas, and Polanyi; WWW 2010)
Although I have not seen a paper combining all the above features in order to predict the helpfulness of a review (or for ranking reviews by helpfulness), I guess that these set of features will bring predictive accuracy pretty close to its limit for this task.

What is next? I guess personalized recommendations are going to appear sooner or later, matching users with reviews that are more likely to benefit them. (Update: See Eugene's comment below for related papers.) For example, a beginner in photography will be interested in a different type of review when buying an SLR, compared to a seasoned professional. We already know that reviews from similar users can be used for recommending products (see Netflix) so it is not unlikely that different types of reviews will be deemed helpful by different types of users.

So, did you find this blog post useful?

Prisoner's Dilemma and Mechanical Turk

I have been reading lately, about the differences between mathematical models of behavior and real human behavior. So, I decided to try on Mechanical Turk the classical game theory model of Prisoner's Dilemma. (See also Brendan's nice explanations and diagrams if you have never been exposed to game theory before.)

From Wikipedia:

In its classical form, the prisoner's dilemma ("PD") is presented as follows:

Two suspects are arrested by the police. The police have insufficient evidence for a conviction, and, having separated both prisoners, visit each of them to offer the same deal. If one testifies (defects from the other) for the prosecution against the other and the other remains silent (cooperates with the other), the betrayer goes free and the silent accomplice receives the full 10-year sentence. If both remain silent, both prisoners are sentenced to only six months in jail for a minor charge. If each betrays the other, each receives a five-year sentence. Each prisoner must choose to betray the other or to remain silent. Each one is assured that the other would not know about the betrayal before the end of the investigation. How should the prisoners act?

If we assume that each player cares only about minimizing his or her own time in jail, then the prisoner's dilemma forms a non-zero-sum game in which two players may each cooperate with or defect from (betray) the other player. In this game, as in all game theory, the only concern of each individual player (prisoner) is maximizing his or her own payoff, without any concern for the other player's payoff. The unique equilibrium for this game is a Pareto-suboptimal solution, that is, rational choice leads the two players to both play defect, even though each player's individual reward would be greater if they both played cooperatively.

My first attempt was to post to Mechanical Turk this dilemma in a setting of the following game:

You are playing a game together with a stranger. Each of you have two choices to play: "trust" or "cheat".
  • If both of you play "trust", you win $30,000 each.
  • If both of you play "cheat", you get $10,000 each.
  • If one player plays "trust" and the other plays "cheat", then the player that played "cheat" gets $50,000 and the player that played "trust" gets $0.
You cannot communicate during the game, and CANNOT see the final action of the other player. Both actions will be revealed simultaneously.

What would you play? "Cheat" or "Trust"?

Basic game theory predicts that the participants will choose "cheat" resulting in a suboptimal equilibrium. However, participants on Mechanical Turk did not behave like that. Instead, 48 out of the 100 participants decided to play "trust", which is above the 33% observed in the lab experiments of (Shafir and Tversky, 1992).

Next, I wanted to make the experiment more realistic. Would anything change if instead of playing an imaginary game, I promised actual monetary benefits to the participants? So, I modified the game, and asked the participants to play against each other. Here is the revised task description.

You are playing a game against another Turker. Your action here will be matched with an action of another Mechanical Turk worker.

Each of you have two choices to play: "trust" or "cheat".
  • If both of you play "trust", you both get a bonus of $0.30.
  • If both of you play "cheat", you both get a bonus of $0.10.
  • If one Turker plays "trust" and the other plays "cheat", then the Turker that played "cheat" gets a bonus of $0.50 and the Turker that played "trust" gets nothing.
What is your action? "Cheat" or "Trust"?

I asked 120 participants to play the game, paying just 1 cent for the participation. Interestingly enough, I had a perfect split in the results. 60 Turkers decided to cheat, and 60 Turkers decided to cheat. The final result was 20 pairs of trust-trust, 20 pairs of cheat-cheat, and 20 pairs of cheat-trust.

In other words, the theory prediction that people will be locked in a non-optimal equilibrium was not correct, neither in the "imaginary" game, nor in the case where the workers had to gain some actually monetary benefit.

Finally, I decided to change the payoff matrix, and replicate the structure of the TV game show "Friend or Foe". There, participants get $50K each if they cooperate, $0 if they do not, and if one chooses trust and the other cheat, the "cheat" gets $100K and the "trust" gets $0.

You are playing a game together with a stranger.

Each of you have two choices to play: "trust" or "cheat".
  • If both of you play "trust", you both win $50,000.
  • If both of you play "cheat", you both get $0.
  • If one player plays "trust" and the other plays "cheat", then the player that played "cheat" gets $100,000 and the player that played "trust" gets $0.
You cannot communicate during the game, and CANNOT see the final action of the other player. Both actions will be revealed simultaneously.

What would you play? "Cheat" or "Trust"?

Interestingly enough, in this setting ALL 100 players ended up playing "trust", which was quite different from the previous game and from the behavior of the players in the TV show, where, in almost 25% of the played games, both players chose "cheat" ending up with $0, and in 25% of the games the players collaborated and played "trust" getting $50K each.

So, in my final attempt, I asked Turkers to play this "Friend of Foe" game, having monetary incentives. Here is the task that I posted on Mechanical Turk.

You are playing a game against another Turker. Your action here will be matched with an action of another Mechanical Turk worker.

Each of you have two choices to play: "trust" or "cheat".
  • If both of you play "trust", you both get a bonus of $0.50.
  • If both of you play "cheat", you both get $0.
  • If one Turker plays "trust" and the other plays "cheat", then the Turker that played "cheat" gets a bonus of $1.00 and the Turker that played "trust" gets nothing.
What is your action? "Cheat" or "Trust"?

In this game, 33% of the users decided to cheat, resulting in 6/50 games where both players got nothing, 23/50 games where both players got a 50 cent bonus, and 21/50 games where one player got $1 and the other player got nothing.

I found the difference in behavior between the imaginary game and the actual one to be pretty interesting. Also, the deviation from the predictions of the game-theoretic model is striking.

Although I am not the first to actually observe that, this deviation got me wondering: Why do we use elaborate game theory models for modeling user behavior, when not even the simplest such models do not correspond to reality? How can someone take seriously the concept of an equilibrium when a game, introduced in the intro chapter of every game theory textbook, simply does not correspond to reality? Do we really understand the limitations of our tools, or mathematical and analytic elegance end up being more important than reality?

 
Free Host | lasik surgery new york