Friday, November 7, 2008

Week 9 Muddiest Point

I was more clear on why XML is a better markup language after our lecture. However, I am still unclear on if it is necessary to create it from scratch as we did in the example--don't most people use a tool to generate it?

Week 10 Readings - The Deep Web

Clearly the theme this week is searching for content on the web! This article confirmed what I think we all instinctively know: there is a great deal of high quality information available that exists in formats that most search engines cannot index. I was particularly intrigued to read that most of Bergman's selected sites in what he calls the Deep Web for his survey are open sites. Yet a typical search done on a search engine most likely will not find it, and this is the way most people look for information on the Web. Many of these sites I've been to myself-Pubmed, NOAA, InfoUSA-and when I think about it, often I just went to the site directly for additional research AFTER using a search engine. I'm sure many in our class will have had that same experience. This really gives me a lot to think about for the role of libraries in providing the "directed query technology" that Bergman advocates. We clearly have a role to play.

Week 10 Readings - Shreves Article

After finishing this article it occurred to me that the relationship between data providers and service providers using the Open Archives Initiative Protocol for Metadata Harvesting has the same pattern as that of search engine companies and every day Internet users. Both are trying to find the right way to index content and provide tools for users to find the appropriate content. I think there is one important difference though; the OAI community relationship are clearly mutually desired and collaborative whereas the search engine to user relationship is desired but not usually collaborative and sometimes adversarial. In the search engine world, search engine providers are often fighting off content providers whose sites clearly won't meet most search queries. OAI providers are working to a "customized
" set of standards to most closely meet their constituents' needs. The search engine providers are operating in a more or less standard less environment--they may have the tougher job!

Week 10 - Web Search Engines Part 1 and Part 2

I thought this two part article gave a mostly easy to understand overview of how search engines work. There were a few highly technical concepts in part 2 that I'm not quite clear on (for example, it wasn't clear to me how the engines identify what a term is) but I was able to understand the basics of the hardware needed and database design required for search engine creation and optimal performance.

Clearly good redundant hardware configuration as well as index design are critical to efficient performance for a search engine. I had no idea that these servers numbered in the hundreds of thousands. After reading through the issues encountered in indexing - spamming and cloaking attempts, dead links, outdated information, secured information - I can understand why no one attempts to index all the content. It's absolutely incredible that these search engines work as well as they do given the formidable challenges that exist in organizing the information.

Saturday, October 18, 2008

Assignment 5 - Koha Virtual Shelf

http://pitt5.kohawc.liblime.com/cgi-bin/koha/bookshelves/shelves.pl?viewshelf=28


For some reason I have two blank entries. They stay even if I check them and click remove. I had deleted some from the bibliographic record after I added them to my shelf, maybe that's why.