Google Scholar now Supports Email Alerts

While searching Google Scholar, I noticed a new icon:


By clicking this button, you can create an email alert, which notifies you for any new papers that may appear for the given query. You can use it:

  • For normal queries, getting notifications for queries such as [author:lastname] or [intitle:titleword], or by any other query on Google Scholar
  • For citation queries, getting information about new papers that cite a given paper. For that, you need to first go to the "Cited by X" page, and then click the alert icon.
At the end, you get a nice list of alerts that can notify you when new papers of a particular author appear on Google Scholar, when new papers about a topic get indexed, or when a paper gets cited (the almighty citation!). Here is how the alert list looks like:



I tries to find some official announcement for this feature and I could not find anything. Not sure if this is rolled out to everyone but it is certainly very very useful!

My Citation Tracker tool becomes less useful now, but I am glad that we will stop trying to implement this feature on top of Google Scholar and Google will start doing it natively. Next step Google: RSS alerts!

Crowdsourcing: Not just a cost saver

Whenever I am discussing crowdsourcing, people tend to ask about the "killer app" for crowdsourcing. In my mind, a new technology is being used first for cost reduction, then due to the afforded flexibility, and finally to achieve tasks that were previously infeasible.

As with many technologies (e.g., VoIP), the first applications of crowdsourcing tend to focus on cost reduction. Label images, moderate comments, verify addresses, collect emails... All tasks that were traditionally done by low-paid interns, are now crowdsourced, for significant cost savings. Although these applications tend to be useful, it is hard to see applications focused on cost cutting to be the ones that will drive the wide adoption of crowdsourcing. In any case, if someone has big tasks like that, they can always contact an outsourcing shop in a developing country, and achieve similar rates.

So, flexibility of deployment tends to be another angle that tends to drive the adoption of crowdsourcing. Pretty much as in the case of cloud computing, crowdsourcing can come pretty handy for companies that need significant resources for just a small period of time. LiveOps was a pioneer of this model, applying "crowdsourcing" before the terms was invented, to dynamically staff call centers for telemarketing companies, handling sudden spikes in demand, and so on.

However, the really interesting applications are those that are truly being enabled by such technology. One such idea was given to me by Amanda Michel of Propublica, last September after a Mechanical Turk Meetup: Propublica wanted to monitor the spending of stimulus money, and examine if the projects assigned to different contractors were being completed properly. By analyzing the collected data, it would be possible to identify systematic problems or outright corruption.

Of course, collecting such data would have been impossible for a journalist, or even for a team of journalists. However, using crowdsourcing it would be pretty easy for local residents to check if a bridge has been fixed, if a road was properly paved, and so on. I loved the idea! It was a clear demonstration of how to take an infeasible task, and use crowdsourcing to make it feasible. Even better: it was encouraging the involvement of citizens and their active collaboration with the government. A win-win project!

Interestingly enough, a recent article on BusinessWeek, indicates that the Obama administration is also employing crowdsourcing for the same task! From the article: "It's the surest way to prevent, say, a convicted contractor from reincorporating a new company under his wife's name and applying for stimulus money, explains Earl Devaney, the special inspector general who oversees stimulus spending. "Only local folks can connect those kinds of dots," he says.


Indeed! If only we had such ideas in Greece...

KDD Accepted Papers, Deadlines for HCOMP 2010 and SNAKDD 2010 workshops

The accepted papers for KDD 2010 are now posted and available at http://www.kdd.org/kdd2010/papers.shtml

As a reminder, the KDD this year will take place in Washington DC, from July 25th to July 28th.

I would also like to draw your attention to two KDD workshops that I am involved with, the Human Computation Workshop (HCOMP 2010) and the Workshop on Social Network Mining and Analysis (SNAKDD 2010). Both have submission deadlines on May 7th. It is pretty easy to travel to DC, so if you have any cool idea, or demo that would be appropriate for these workshops, please submit!

For those too bored to visit the respective websites, here are the call for papers:

Human Computation Workshop (HCOMP 2010) - Call for Papers

Most research in data mining and knowledge discovery relies heavily on the availability of datasets. With the rapid growth of user generated content on the internet, there is now an abundance of sources from which data can be drawn. Compared to the amount of work in the field on techniques for pattern discovery and knowledge extraction, there has been little effort directed at the study of effective methods for collecting and evaluating the quality of data.

Human computation is a relatively new research area that studies the process of channeling the vast internet population to perform tasks or provide data towards solving difficult problems that no known efficient computer algorithms can yet solve. There are various genres of human computation applications available today. Games with a purpose (e.g., the ESP Game) specifically target online gamers who, in the process of playing an enjoyable game, generate useful data (e.g., image tags). Crowdsourcing marketplaces (e.g. Amazon Mechanical Turk) are human computation applications that coordinate workers to perform tasks in exchange for monetary rewards. In identity verification tasks, users need to perform some computation in order to access some online content; one example of such a human computation application is reCAPTCHA, which leverages millions of users who solve CAPTCHAs every day to correct words in books that optical character recognition (OCR) programs fail to recognize with certainty.

Human computation is an area with significant research challenges and increasing business interest, making this doubly relevant to KDD. KDD provides an ideal forum for a workshop on human computation as a form of cost-sensitive data acquisition. The workshop also offers a chance to bring in practitioners with complementary real-world expertise in gaming and mechanism design who might not otherwise attend this academic conference.

The first Human Computation Workshop (HComp 2009) was held on June 28th, 2009, in Paris, France, collocated with KDD 2009. The overall themes that emerged from this workshop were very clear: on the one hand, there is the experimental side of human computation, with research on new incentives for users to participate, new types of actions, and new modes of interaction. This includes work on new programming paradigms and game templates designed to enable rapid prototyping, allow partial completion of tasks, and aid in reusability of game design. On the more theoretic side, we have research modeling these actions and incentives to examine what theory predicts about these designs. Finally, there is work on noisy results generated by such games and systems: how can we best handle noise, identify labeler expertise, and use the generated data for data mining purposes?

Learning from HComp 2009, we have expanded the topics of relevance to the workshop. The goal of HComp 2010 is to bring together academic and industry researchers in a stimulating discussion of existing human computation applications and future directions of this new subject area. We solicit papers related to various aspects of both general human computation techniques and specific applications, e.g. general design principles; implementation; cost-benefit analysis; theoretical approaches; privacy and security concerns; and incorporation of machine learning / artificial intelligence techniques. An integral part of this workshop will be a demo session where participants can showcase their human computation applications. Specifically, topics of interests include, but are not limited to:

  • Abstraction of human computation tasks into taxonomies of mechanisms
  • Theories about what makes some human computation tasks fun and addictive
  • Differences between collaborative vs. competitive tasks
  • Programming languages, tools and platforms to support human computation
  • Domain-specific implementation challenges in human computation games
  • Cost, reliability, and skill of labelers
  • Benefits of one-time versus repeated labeling
  • Game-theoretic mechanism design of incentives for motivation and honest reporting
  • Design of manipulation-resistance mechanisms in human computation
  • Effectiveness of CAPTCHAs
  • Concerns regarding the protection of labeler identities
  • Active learning from imperfect human labelers
  • Creation of intelligent bots in human computation games
  • Utility of social networks and social credit in garnering data
  • Optimality in the context of human computation
  • Focus on tasks where crowds, not individuals, have the answers
  • Limitations of human computation


Workshop on Social Network Mining and Analysis - Call for Papers

Social networks research has come a long way since the notable �six-degree separation� experiment. In recent years, social network research has advanced significantly, thanks to the prevalence of the online social websites and the availability of a variety of offline large-scale social network systems such as collaboration networks. These social network systems are usually characterized by the complex network structures and rich accompanying contextual information. Researchers are increasingly interested in addressing a wide range of challenges residing in these disparate social network systems, including identifying common static topological properties and dynamic properties during the formation and evolution of these social networks, and how contextual information can help in analyzing the pertaining social networks. These issues have important implications on community discovery, anomaly detection, trend prediction and can enhance applications in multiple domains such as information retrieval, recommendation systems, security and so on.

The fourth SNA-KDD '2010 aims to bring together practitioners and researchers with a specific focus on the emerging trends and industry needs associated with the traditional Web, the social Web, and other forms of social networking systems. Both theoretical and experimental submissions are encouraged. The interesting topics include (1) data mining advances on the discovery and analysis of communities, on personalization for solitary activities (like search) and social activities (like discovery of potential friends), on the analysis of user behavior in open fora (like conventional sites, blogs and fora) and in commercial platforms (like e-auctions) and on the associated security and privacy-preservation challenges; (2) social network modeling, scalable, customizable social network infrastructure construction, dynamic growth and evolution patterns identification and discovery using machine learning approaches or multi-agent based simulation.

The fourth SNA-KDD '2010 solicits contributions on social network analysis and graph mining, including the emerging applications of the Web as a social medium. Papers should elaborate on data mining methods, issues associated to data preparation and pattern interpretation, both for conventional data (usage logs, query logs, document collections) and for multimedia data (pictures and their annotations, multi-channel usage data). Topics of interest include but are not limited to:

  • Communities discovery and analysis in large scale online and offline social networks
  • Personalization for search and for social interaction
  • Recommendations for product purchase, information acquisition and establishment of social relations
  • Data protection inside communities
  • Misbehavior detection in communities
  • Web mining algorithms for clickstreams, documents and search streams
  • Preparing data for web mining
  • Pattern presentation for end-users and experts
  • Evolution of patterns in the Web
  • Evolution of communities in the Web
  • Dynamics and evolution patterns of social networks, trend prediction
  • Contextual social network analysis
  • Temporal analysis on social networks topologies
  • Search algorithms on social networks
  • Multi-agent based social network modeling and analysis
  • Application of social network analysis
  • Anomaly detection in social network evolution

Yahoo!'s Key Scientific Challenges: Your student is a winner!

I got the following email:

Dear Panos,

The judging is done. From an outstanding group of 200 proposals, twenty-two exceptional PhD students have been selected to be part of Yahoo!�s 2010 Key Scientific Challenges (KSC) Program. Your student, Nikolay Archak has been chosen to receive this very competitive award. Congratulations!

Supporting the academic community is a top priority at Yahoo!. We created the KSC Program to support a limited number of outstanding PhD students who we believe are doing research in very important and challenging areas. The Program provides each student with $5,000 of unrestricted funds for the support of their research activities (e.g., conference fees and travel, lab materials, professional society membership dues, etc.). The funds are distributed through the university and paid directly to the student for use at their discretion. As part of this program, the student also receives an exclusive invitation to a unique workshop (most likely to be held in August at our location in Sunnyvale, California), where we will focus on novel disciplines and important technical challenges for the Internet research community, slanted specifically for graduate students whose innovative work is just emerging. We think that the opportunity to interact with Yahoo! scientists and other top graduate students in an informal and supportive environment without having to submit papers for review will provide a unique forum for open and stimulating discussion of work in its early stages.

Students will also have the opportunity to work with select datasets through our Webscope program. Yahoo! will cover all travel expenses to the KSC Workshop independent of the $5,000 award.

We are excited to be able to support these students with KSC grants, and we look forward to having them as part of our broad research alliance of outstanding students and faculty. Please feel free to contact Jamie Lockwood at jamieloc@yahoo-inc.com if you have any questions regarding our KSC Program.

I offer you my warmest congratulations. We eagerly look forward to hosting Nikolay at our KSC Workshop as part of the Yahoo! family.

Best regards,

Ken Schmidt
Director, Academic Relations

 Nikolay is not a stranger to winning competitions. He has won in the past 3 times the TopCoder competition, ended up 2nd two more times. He also has a streak of accepted publications, including a single-authored paper at WWW2010 this year.

Nikolay, congratulations!

Stop Publishing!

The last few months, I feel that I have an endless queue of reviewing tasks to complete. WWW, followed by DBRank, followed by EC, followed by KDD, followed by VLDB, followed by WebDB, plus an NSF panel, plus some journal reviews, and I have rejected invitations for a few additional conferences including SIGIR, SIGMOD, and a few others. This puts my count at least 40 reviews over the last 4-5 months. (Just to break even, I will need to submit 10-12 papers.)

Needless to say, having such a reviewing load means that I cannot really do a good job in reviewing. My reviews have been declining in quality, signalling that I need to learn to say no.

At the same time, I also notice that the other reviews that are being submitted are not that great either. While on the one hand I feel happy ("OK, I am not that bad"), on the other hand I feel that this cannot be good. If nobody has time to review thoroughly, what is the whole point of peer reviewing?

One solution is to accept fewer invitations for PCs, allowing for more time per paper. However, I know that without volunteering time and effort for reviewing the system cannot work! There are simply not that many reviewers available!

Part of this problem is, of course, the increased need to get papers published: For tenure, for getting a job, even for being admitted to a PhD program! I feel that there is something wrong when, to be admitted to a PhD program, you need to have already research experience. This increased need for more and more publications, overloads the reviewing system which unfortunately has a limited capacity.

Unfortunately, it is not easy to reverse this trend. The incentives are setup in a way to encourage quantity of publications, preferably in good venues. Once the paper gets accepted in a good venue, the goal is achieved. This encourages publications that are "good enough" to pass the reviewing process, not papers that have stellar quality. And with the increased noise in the reviewing process, the distinction between "good enough to be published" and "what the hell, send it, we may get lucky" is getting blurrier and blurrier. In fact, I have cases in my papers that the reviewers did such a poor job that I never understood at the end if my paper was worth getting published, or I got just lucky.

I noticed though a positive development! Through the Greek University Reform Forum, I learned that:


The German Research Society (DFG) has introduced new guidelines for applications and evaluations of proposals, which will be valid as of July 1, 2010.

A rough-and-ready translation of the main points:
  • Applicants should cite in their CV only up to FIVE publications, those which are most relevant for the proposal at hand;
  • In reports about running projects, a maximum of TWO publications PER YEAR. In case of projects with more than one PIs, a maximum of THREE publications PER YEAR.
  • The goal of the new guideleines is to put emphasis on quality instead of quantity and to stop the flood of publishing for the sake of the numbers.

It has caused quite some stir here in Germany and the voices to enforce such rules also for decisions on faculty positions are getting louder.

While some of these ideas are already in place (e.g., NSF also allows only five publications in the CV), the idea of "counting" only two publications per year for each project is definitely a step towards the right direction. It is not going to be trivial to reverse the "get as many publications in top venues as possible" trend, but every step towards de-emphasizing quantity counts.

After all, Pollock was also getting paid by the piece when he worked for the Federal Art Project, but none of his famous paintings come from that period.

 
Free Host | lasik surgery new york