Search   Feed   Browse   Add
Feed items 1 - 5 of 5 for August 2008

Clever method of near duplicate detection - August 7, 2008

Martin Theobald, Jonathan Siddharth, and Andreas Paepcke from Stanford University have a cute idea in their SIGIR 2008 paper, "SpotSigs: Robust and Efficient Near Duplicate Detection in Large Web Collections" (PDF). They focus near duplicate detection on the important parts of a web pages by using the next few words after a stop word as a signature.An extended excerpt:The frequent presence of many diverse semantic units in individual Web pages makes near-duplicate detection particularly...
http://glinden.blogspot.com/2008/08/clever-method-of-near-duplicate.html

BrowseRank: Ranking pages by how people use them - August 6, 2008

Liu et al. from Microsoft Research Asia had the best student paper at SIGIR 2008, "BrowseRank: Letting Web Users Vote for Page Importance" (PDF), that builds a "user browsing graph" of web pages where "edges represent real transitions between web pages by users."An excerpt:The user browsing graph can more precisely represent the web surfer's random walk process, and thus is more useful for calculating page importance. The more visits of the page made by users and the longer time periods spent..
http://glinden.blogspot.com/2008/08/browserank-ranking-pages-by-how-people.html

Caching, index pruning, and the query stream - August 5, 2008

A SIGIR 2008 paper out of Yahoo Research, "ResIn: A Combination of Results Caching and Index Pruning for High-performance Web Search Engines" (ACM page), looks at how performance optimizations to a search engine can impact each other. In particular, it looks at how caching the results of search queries impacts the query load that needs to be served from the search index, and therefore changes the effectiveness of index pruning, which attempts to serve some queries out of a reduced index.An...
http://glinden.blogspot.com/2008/08/caching-index-pruning-and-query-stream.html

To personalize or not to personalize - August 4, 2008

Jaime Teevan, Sue Dumais, and Dan Liebling had a paper at SIGIR 2008, "To Personalize or Not to Personalize: Modeling Queries with Variation in User Intent" (PDF), that looks at how and when different people want different things from the same search query.An excerpt:For some queries, everyone ... is looking for the same thing. For other queries, different people want very different results.We characterize queries using features of the query, the results returned for the query, and people's...
http://glinden.blogspot.com/2008/08/to-personalize-or-not-to-personalize.html

Modeling how searchers look at search results - August 1, 2008

Georges Dupret and Benjamin Piwowarski from Yahoo Research had a great paper at SIGIR 2008, "A User Browsing Model to Predict Search Engine Click Data from Past Observations" (ACM page). It nicely extends earlier models of searcher behavior to allow for people skipping results after hitting a lot of uninteresting results.An extended excerpt:Click data seems the perfect source of information when deciding which documents (or ads) to show in answer to a query. It can be thought of as the...
http://glinden.blogspot.com/2008/08/modeling-how-searchers-look-at-search.html
Available Archives
- June (5 items)
- July (15 items)
- August (5 items)
Sponsored Links
© 2008 FeedCapsule.com  |  Contact